Group
Files and media
Image, video, audio, document, screenshot, and upload capabilities.
Image, video, audio, document, screenshot, and upload capabilities.
The index groups these capabilities by job rather than vendor.
- Audio file upload and understanding
Upload recorded audio and expose its supported speech, timing, speaker, or acoustic content.
- File and media inputs
Track accepted file types, model-visible content, preprocessing, and upload limits separately.
- File upload limits and quota visibility
Record numeric attachment size, count, page, duration, storage, and rolling quotas.
- Image upload and paste
Attach, paste, or select an image whose visual content reaches the model.
- PDF documents
Upload a PDF and make its text, layout, and supported visual content available as context.
- Realtime voice
Speak and listen over a live audio session.
- Screenshots
Capture the current screen or window as context.
- Text and office document input
Upload common text, word-processing, presentation, and spreadsheet formats as model context.
- Video preprocessing disclosure
Explain frame, audio, transcript, resizing, truncation, and context-accounting behavior.
- Video upload and understanding
Upload a video and expose documented visual, audio, transcript, and timing information to the model.