Product feature · perception
Audio file inputCommon product term
Upload recorded audio for use as model input.
Terminology basis: Common product term. xAI — Grok files and data FAQ
Audio file input: 5 supported, 0 partial, 0 unsupported, 26 unreviewed across 31 cataloged products.
Current evidence by product
Can my agent use Audio file input?
Read across for the answer. 5 of 31 current product columns have reviewed evidence; unreviewed does not mean unsupported.
Web
9 productsDesktop
13 productsCLI
9 productsDefinition and scope
What this capability means
This row covers recorded audio supplied as a file, not a live duplex voice session. Evidence should distinguish speech transcription from speaker diarization, timestamps, language identification, non-speech sound analysis, emotion or prosody claims, and native audio-model input.
Record formats, channels, maximum bytes and duration, language limits, transcript editability, retention, model and plan requirements, and whether tools or sub-agents receive the original recording or only a derivative transcript.
Traceable compatibility
Assertion ledger
Documentation evidence only. No runtime conformance test is implied.
- Target
- 2026-08-28 Grok Web documentation observation · hosted-observation
- Environment
- hosted-default
- Observed
- 2026-08-28
- runtimetranscription and interpretation are documented, but accepted formats, diarization, timestamp, and native-audio fidelity are not specified on the reviewed page
- xAI — Grok files and data FAQdocumented · reviewed 2026-08-28
- Target
- 2026-08-28 Grok Bot desktop documentation observation · hosted-observation
- Environment
- hosted-default
- Observed
- 2026-08-28
- runtimeaudio is a supported attachment up to 25 MB, but transcript, speaker, timing, and acoustic semantics are not documented on the reviewed page
- xAI — Grok Bot files and resultsdocumented · reviewed 2026-08-28
- Target
- 2026-08-28 Gemini Apps documentation observation · hosted-observation
- Environment
- hosted-default
- Observed
- 2026-08-28
- runtimedirect audio upload and analysis are documented
- plantotal audio is limited to 10 minutes without a Google AI plan and 3 hours with Google AI Pro or Ultra
- runtimeaccepted audio formats, diarization, timestamps, transcript fidelity, and native acoustic semantics are not established by the reviewed page
- Google Gemini Apps Help — Upload and analyze filesdocumented · reviewed 2026-08-28
- Target
- 2026-08-28 Perplexity web documentation observation · hosted-observation
- Environment
- hosted-default
- Observed
- 2026-08-28
- runtimedirect audio uploads are automatically transcribed; the documented formats are MPEG, WAV, AIFF, OGG, FLAC, and MP3, with a 40 MB general per-file limit
- runtimespeaker identification and labels are documented, but timestamps and non-speech acoustic or prosody fidelity are not established
- Perplexity Help Center — File uploadsdocumented · reviewed 2026-08-28
- Target
- current Gemini CLI custom-command documentation · dated-documentation
- Environment
- local-default
- Observed
- 2026-08-28
- runtimesupported audio files referenced with @{...} inside a custom command are encoded and injected as multimodal input
- runtimeaccepted audio formats, ordinary-prompt attachment methods, transcripts, timestamps, diarization, and acoustic semantics are not established by the reviewed page
- Gemini CLI — Custom commandsdocumented · reviewed 2026-08-28
- 1. Evidence checked 2026-08-28: xAI documents direct audio uploads in Grok chats and describes transcription and interpretation of audio and video inputs; detailed formats, timing, diarization, and duration limits are not stated on the reviewed FAQ page.
- 2. Evidence checked 2026-08-28: Grok Bot lists audio among common supported inputs and documents a 25 MB per-audio-file limit, but the reviewed page does not specify transcript, speaker, timing, or acoustic fidelity.
- 3. Evidence checked 2026-08-28: Gemini Apps accept audio uploads and document total audio-duration limits of 10 minutes without a Google AI plan or 3 hours with Google AI Pro or Ultra. The reviewed page does not define diarization, timestamp, transcript, or acoustic-understanding fidelity.
- 4. Evidence checked 2026-08-28: Perplexity web file uploads accept MPEG, WAV, AIFF, OGG, FLAC, and MP3 audio, automatically transcribe spoken content, and can identify and label speakers. The general file-upload page documents a 40 MB per-file limit.
- 5. Evidence checked 2026-08-28: Gemini CLI custom commands encode a supported audio path referenced with @{...} and inject it as multimodal input. The reviewed page does not enumerate audio formats or establish transcript, timestamp, diarization, or acoustic-analysis fidelity.
No sourced issues are attached to this row.
- Methodology note
- xAI — Grok files and data FAQ documented · xAI · reviewed 2026-08-28
- xAI — Grok Bot files and results documented · xAI · reviewed 2026-08-28
- Google Gemini Apps Help — Upload and analyze files documented · Google · reviewed 2026-08-28
- Perplexity Help Center — File uploads documented · Perplexity · reviewed 2026-08-28
- Gemini CLI — Custom commands documented · Google · reviewed 2026-08-28
Found a wrong or incomplete result?
Name the exact product, describe what happened, add the date, and link public evidence when available. A community report starts a review. It does not change the support state by itself.