---
title: "Audio file input"
canonical: "https://canmyagentuse.com/features/audio-file-input"
contentKind: "feature"
locale: "en"
description: "Upload recorded audio for use as model input."
llmSummary: "Audio file input covers recorded uploads and is separate from realtime voice. Transcription, timestamps, speaker labels, and acoustic analysis are recorded as qualifiers only when documented."
publishedAt: "2026-08-28T00:00:00.000Z"
updatedAt: "2026-08-28T00:00:00.000Z"
verifiedAt: "2026-08-28"
tags: ["perception","audio","uploads","transcription"]
---

# Audio file input

Audio file input covers recorded uploads and is separate from realtime voice. Transcription, timestamps, speaker labels, and acoustic analysis are recorded as qualifiers only when documented.

- HTML: https://canmyagentuse.com/features/audio-file-input
- JSON: https://canmyagentuse.com/api/v1/features/audio-file-input.json
- Markdown: https://canmyagentuse.com/features/audio-file-input.md

Terminology basis: **Common product term** — https://docs.x.ai/grok/faq.

## Current support at a glance

Audio file input: 5 supported, 0 partial, 0 unsupported, 26 unreviewed across 31 cataloged products.

- Reviewed current products: 5 of 31
- Supported: 5
- Partial: 0
- Unsupported: 0
- Unreviewed: 26
- Not applicable: 0

Unknown or unreviewed means insufficient published evidence; it does not mean unsupported.

This row covers recorded audio supplied as a file, not a live duplex voice session. Evidence should distinguish speech transcription from speaker diarization, timestamps, language identification, non-speech sound analysis, emotion or prosody claims, and native audio-model input.

Record formats, channels, maximum bytes and duration, language limits, transcript editability, retention, model and plan requirements, and whether tools or sub-agents receive the original recording or only a derivative transcript.

## Catalog context

- Category: [perception](/categories/perception.md)
- Terminology basis: Common product term
- Aliases: audio input, audio upload, recording attachment, speech transcription
- Family: [File and media inputs](/features/file-inputs.md)
- Siblings: [Document input](/features/office-document-input.md), [File upload limits](/features/upload-limits.md), [Image input](/features/image-input.md), [PDF input](/features/pdf-documents.md), [Video input](/features/video-input.md)

## Compatibility assertions

Unknown means insufficient published evidence; it does not mean unsupported.

### ChatGPT (web)

- Harness: [ChatGPT](/harnesses/chatgpt-web.md)
- current: **Unknown**
- preview: **Unknown**

### Claude (web)

- Harness: [Claude](/harnesses/claude-web.md)
- current: **Unknown**
- preview: **Unknown**

### Gemini (web)

- Harness: [Gemini](/harnesses/gemini-web.md)
- current: **Supported**
  - Target: hosted-observation — 2026-08-28 Gemini Apps documentation observation; observed 2026-08-28
  - Environment: hosted-default
  - Constraint (runtime): direct audio upload and analysis are documented
  - Constraint (plan): total audio is limited to 10 minutes without a Google AI plan and 3 hours with Google AI Pro or Ultra
  - Constraint (runtime): accepted audio formats, diarization, timestamps, transcript fidelity, and native acoustic semantics are not established by the reviewed page
  - Evidence: [Google Gemini Apps Help — Upload and analyze files](https://support.google.com/gemini/answer/14903178?co=GENIE.Platform%3DDesktop&hl=en) — documented; observed 2026-08-28
  - Qualification note 3: Evidence checked 2026-08-28: Gemini Apps accept audio uploads and document total audio-duration limits of 10 minutes without a Google AI plan or 3 hours with Google AI Pro or Ultra. The reviewed page does not define diarization, timestamp, transcript, or acoustic-understanding fidelity.
- preview: **Unknown**

### Copilot (web)

- Harness: [Copilot](/harnesses/copilot-web.md)
- current: **Unknown**
- preview: **Unknown**

### Grok (web)

- Harness: [Grok](/harnesses/grok-web.md)
- current: **Supported**
  - Target: hosted-observation — 2026-08-28 Grok Web documentation observation; observed 2026-08-28
  - Environment: hosted-default
  - Constraint (runtime): transcription and interpretation are documented, but accepted formats, diarization, timestamp, and native-audio fidelity are not specified on the reviewed page
  - Evidence: [xAI — Grok files and data FAQ](https://docs.x.ai/grok/faq) — documented; observed 2026-08-28
  - Qualification note 1: Evidence checked 2026-08-28: xAI documents direct audio uploads in Grok chats and describes transcription and interpretation of audio and video inputs; detailed formats, timing, diarization, and duration limits are not stated on the reviewed FAQ page.
- preview: **Unknown**

### Grok Bot (desktop)

- Harness: [Grok Bot](/harnesses/grok-bot-desktop.md)
- current: **Supported**
  - Target: hosted-observation — 2026-08-28 Grok Bot desktop documentation observation; observed 2026-08-28
  - Environment: hosted-default
  - Constraint (runtime): audio is a supported attachment up to 25 MB, but transcript, speaker, timing, and acoustic semantics are not documented on the reviewed page
  - Evidence: [xAI — Grok Bot files and results](https://docs.x.ai/grok-bot/files-and-results) — documented; observed 2026-08-28
  - Qualification note 2: Evidence checked 2026-08-28: Grok Bot lists audio among common supported inputs and documents a 25 MB per-audio-file limit, but the reviewed page does not specify transcript, speaker, timing, or acoustic fidelity.

### Perplexity (web)

- Harness: [Perplexity](/harnesses/perplexity-web.md)
- current: **Supported**
  - Target: hosted-observation — 2026-08-28 Perplexity web documentation observation; observed 2026-08-28
  - Environment: hosted-default
  - Constraint (runtime): direct audio uploads are automatically transcribed; the documented formats are MPEG, WAV, AIFF, OGG, FLAC, and MP3, with a 40 MB general per-file limit
  - Constraint (runtime): speaker identification and labels are documented, but timestamps and non-speech acoustic or prosody fidelity are not established
  - Evidence: [Perplexity Help Center — File uploads](https://www.perplexity.ai/help-center/en/articles/10354807-file-uploads) — documented; observed 2026-08-28
  - Qualification note 4: Evidence checked 2026-08-28: Perplexity web file uploads accept MPEG, WAV, AIFF, OGG, FLAC, and MP3 audio, automatically transcribe spoken content, and can identify and label speakers. The general file-upload page documents a 40 MB per-file limit.
- preview: **Unknown**

### Le Chat (web)

- Harness: [Le Chat](/harnesses/le-chat.md)
- current: **Unknown**
- preview: **Unknown**

### Devin (web)

- Harness: [Devin](/harnesses/devin-web.md)
- current: **Unknown**
- preview: **Unknown**

### Replit Agent (web)

- Harness: [Replit Agent](/harnesses/replit-agent.md)
- current: **Unknown**
- preview: **Unknown**

### ChatGPT (desktop)

- Harness: [ChatGPT](/harnesses/chatgpt-desktop.md)
- current: **Unknown**
- preview: **Unknown**

### Claude (desktop)

- Harness: [Claude](/harnesses/claude-desktop.md)
- current: **Unknown**
- preview: **Unknown**

### Cursor (desktop)

- Harness: [Cursor](/harnesses/cursor.md)
- current: **Unknown**
- preview: **Unknown**

### OpenWork Desktop (desktop)

- Harness: [OpenWork Desktop](/harnesses/openwork-desktop.md)
- current: **Unknown**

### Copilot Chat (desktop)

- Harness: [Copilot Chat](/harnesses/vscode-copilot.md)
- current: **Unknown**
- preview: **Unknown**

### Chrome WebMCP origin trial (desktop)

- Harness: [Chrome WebMCP origin trial](/harnesses/chrome-webmcp-preview.md)
- current: **Unknown**

### Windsurf (desktop)

- Harness: [Windsurf](/harnesses/windsurf.md)
- current: **Unknown**
- preview: **Unknown**

### Zed Agent (desktop)

- Harness: [Zed Agent](/harnesses/zed-agent.md)
- current: **Unknown**
- preview: **Unknown**

### Continue (desktop)

- Harness: [Continue](/harnesses/continue.md)
- current: **Unknown**
- preview: **Unknown**

### Cline (desktop)

- Harness: [Cline](/harnesses/cline.md)
- current: **Unknown**
- preview: **Unknown**

### JetBrains AI (desktop)

- Harness: [JetBrains AI](/harnesses/jetbrains-ai.md)
- current: **Unknown**
- preview: **Unknown**

### Warp (desktop)

- Harness: [Warp](/harnesses/warp.md)
- current: **Unknown**
- preview: **Unknown**

### Claude CLI (cli)

- Harness: [Claude CLI](/harnesses/claude-cli.md)
- current: **Unknown**
- preview: **Unknown**

### ChatGPT CLI (cli)

- Harness: [ChatGPT CLI](/harnesses/chatgpt-cli.md)
- current: **Unknown**
- preview: **Unknown**

### Codex CLI (cli)

- Harness: [Codex CLI](/harnesses/codex-cli.md)
- current: **Unknown**
- preview: **Unknown**

### OpenCode (cli)

- Harness: [OpenCode](/harnesses/opencode.md)
- current: **Unknown**
- preview: **Unknown**

### Gemini CLI (cli)

- Harness: [Gemini CLI](/harnesses/gemini-cli.md)
- current: **Supported**
  - Target: dated-documentation — current Gemini CLI custom-command documentation; observed 2026-08-28
  - Environment: local-default
  - Constraint (runtime): supported audio files referenced with @{...} inside a custom command are encoded and injected as multimodal input
  - Constraint (runtime): accepted audio formats, ordinary-prompt attachment methods, transcripts, timestamps, diarization, and acoustic semantics are not established by the reviewed page
  - Evidence: [Gemini CLI — Custom commands](https://geminicli.com/docs/cli/custom-commands/) — documented; observed 2026-08-28
  - Qualification note 5: Evidence checked 2026-08-28: Gemini CLI custom commands encode a supported audio path referenced with @{...} and inject it as multimodal input. The reviewed page does not enumerate audio formats or establish transcript, timestamp, diarization, or acoustic-analysis fidelity.
- preview: **Unknown**

### Aider (cli)

- Harness: [Aider](/harnesses/aider.md)
- current: **Unknown**
- preview: **Unknown**

### Goose (cli)

- Harness: [Goose](/harnesses/goose.md)
- current: **Unknown**
- preview: **Unknown**

### Copilot CLI (cli)

- Harness: [Copilot CLI](/harnesses/copilot-cli.md)
- current: **Unknown**
- preview: **Unknown**

### Amp (cli)

- Harness: [Amp](/harnesses/amp-cli.md)
- current: **Unknown**
- preview: **Unknown**
