---
title: "Files and media"
canonical: "https://canmyagentuse.com/categories/perception"
contentKind: "category"
locale: "en"
description: "Image, video, audio, document, screenshot, and upload capabilities."
llmSummary: "Category Files and media: Image, video, audio, document, screenshot, and upload capabilities."
publishedAt: "2026-08-28T00:00:00.000Z"
updatedAt: "2026-08-28T00:00:00.000Z"
verifiedAt: null
tags: ["category"]
---

# Files and media

Category Files and media: Image, video, audio, document, screenshot, and upload capabilities.

- HTML: https://canmyagentuse.com/categories/perception
- Markdown: https://canmyagentuse.com/categories/perception.md

Image, video, audio, document, screenshot, and upload capabilities.

The index groups these capabilities by job rather than vendor.

## Capabilities

- [Audio file input](/features/audio-file-input.md) — Audio file input covers recorded uploads and is separate from realtime voice. Transcription, timestamps, speaker labels, and acoustic analysis are recorded as qualifiers only when documented.
- [Document input](/features/office-document-input.md) — Document input means uploaded text, word-processing, presentation, or spreadsheet content is available as model input. Supported formats and extraction fidelity are qualifiers.
- [File and media inputs](/features/file-inputs.md) — File and media inputs is an internal catalog grouping for image, PDF, office-document, audio, video, and upload-limit questions.
- [File upload limits](/features/upload-limits.md) — File upload limits record documented size, count, page, duration, frequency, and storage limits with their scope. Accepting one test file does not establish every limit.
- [Image input](/features/image-input.md) — Image input means a product accepts attached, pasted, or selected images as model input. Supported methods, formats, limits, and surfaces are recorded as qualifiers.
- [PDF input](/features/pdf-documents.md) — PDF input makes an uploaded PDF available as model input. Text extraction, visual page analysis, file limits, and attachment limits are qualifiers.
- [Realtime voice](/features/realtime-voice.md) — Realtime voice supports live spoken input and audio responses. Turn-taking, interruption, latency, models, and session limits are qualifiers.
- [Screenshots](/features/screenshots.md) — Screenshot support captures a screen, window, or browser page for use as model input. Browser-only and arbitrary desktop capture are distinct qualifiers.
- [Video input](/features/video-input.md) — Video input means an uploaded video contributes model context. Frame, audio, transcript, timing, and format behavior are recorded as qualifiers only when documented.
