Product feature · perception

Audio file inputCommon product term

Upload recorded audio for use as model input.

Terminology basis: Common product term. xAI — Grok files and data FAQ

Audio file input: 5 supported, 0 partial, 0 unsupported, 26 unreviewed across 31 cataloged products.

Markdown · JSON

Explore this familyMore in File and media inputs5 capabilities

Current evidence by product

Can my agent use Audio file input?

Read across for the answer. 5 of 31 current product columns have reviewed evidence; unreviewed does not mean unsupported.

  • Supported5
  • Partial0
  • Unsupported0
  • Unknown26
  • Not applicable0

Web

9 products
GrokGrok
Y1Supported
Observed 2026-08-28Current · 1 condition
?Unknown
No source reviewedPreview record
Report this result

Desktop

13 products

CLI

9 products

Unknown means no public evidence has been reviewed for that product and capability. It does not mean unsupported.

How statuses are assigned

Definition and scope

What this capability means

This row covers recorded audio supplied as a file, not a live duplex voice session. Evidence should distinguish speech transcription from speaker diarization, timestamps, language identification, non-speech sound analysis, emotion or prosody claims, and native audio-model input.

Record formats, channels, maximum bytes and duration, language limits, transcript editability, retention, model and plan requirements, and whether tools or sub-agents receive the original recording or only a derivative transcript.

Traceable compatibility

Assertion ledger

Documentation evidence only. No runtime conformance test is implied.

Grokweb · current
Supported
Target
2026-08-28 Grok Web documentation observation · hosted-observation
Environment
hosted-default
Observed
2026-08-28
  • runtimetranscription and interpretation are documented, but accepted formats, diarization, timestamp, and native-audio fidelity are not specified on the reviewed page
Evidence
Grok Botdesktop · current
Supported
Target
2026-08-28 Grok Bot desktop documentation observation · hosted-observation
Environment
hosted-default
Observed
2026-08-28
  • runtimeaudio is a supported attachment up to 25 MB, but transcript, speaker, timing, and acoustic semantics are not documented on the reviewed page
Evidence
Geminiweb · current
Supported
Target
2026-08-28 Gemini Apps documentation observation · hosted-observation
Environment
hosted-default
Observed
2026-08-28
  • runtimedirect audio upload and analysis are documented
  • plantotal audio is limited to 10 minutes without a Google AI plan and 3 hours with Google AI Pro or Ultra
  • runtimeaccepted audio formats, diarization, timestamps, transcript fidelity, and native acoustic semantics are not established by the reviewed page
Evidence
Perplexityweb · current
Supported
Target
2026-08-28 Perplexity web documentation observation · hosted-observation
Environment
hosted-default
Observed
2026-08-28
  • runtimedirect audio uploads are automatically transcribed; the documented formats are MPEG, WAV, AIFF, OGG, FLAC, and MP3, with a 40 MB general per-file limit
  • runtimespeaker identification and labels are documented, but timestamps and non-speech acoustic or prosody fidelity are not established
Evidence
Gemini CLIcli · current
Supported
Target
current Gemini CLI custom-command documentation · dated-documentation
Environment
local-default
Observed
2026-08-28
  • runtimesupported audio files referenced with @{...} inside a custom command are encoded and injected as multimodal input
  • runtimeaccepted audio formats, ordinary-prompt attachment methods, transcripts, timestamps, diarization, and acoustic semantics are not established by the reviewed page
Evidence
  1. 1. Evidence checked 2026-08-28: xAI documents direct audio uploads in Grok chats and describes transcription and interpretation of audio and video inputs; detailed formats, timing, diarization, and duration limits are not stated on the reviewed FAQ page.
  2. 2. Evidence checked 2026-08-28: Grok Bot lists audio among common supported inputs and documents a 25 MB per-audio-file limit, but the reviewed page does not specify transcript, speaker, timing, or acoustic fidelity.
  3. 3. Evidence checked 2026-08-28: Gemini Apps accept audio uploads and document total audio-duration limits of 10 minutes without a Google AI plan or 3 hours with Google AI Pro or Ultra. The reviewed page does not define diarization, timestamp, transcript, or acoustic-understanding fidelity.
  4. 4. Evidence checked 2026-08-28: Perplexity web file uploads accept MPEG, WAV, AIFF, OGG, FLAC, and MP3 audio, automatically transcribe spoken content, and can identify and label speakers. The general file-upload page documents a 40 MB per-file limit.
  5. 5. Evidence checked 2026-08-28: Gemini CLI custom commands encode a supported audio path referenced with @{...} and inject it as multimodal input. The reviewed page does not enumerate audio formats or establish transcript, timestamp, diarization, or acoustic-analysis fidelity.