Product feature · perception

Video inputCommon product term

Upload or select a video for use as model input.

Terminology basis: Common product term. xAI — Grok files and data FAQ

Video input: 4 supported, 1 partial, 0 unsupported, 26 unreviewed across 31 cataloged products.

Markdown · JSON

Explore this familyMore in File and media inputs5 capabilities

Current evidence by product

Can my agent use Video input?

Read across for the answer. 5 of 31 current product columns have reviewed evidence; unreviewed does not mean unsupported.

  • Supported4
  • Partial1
  • Unsupported0
  • Unknown26
  • Not applicable0

Web

9 products
GrokGrok
Y1Supported
Observed 2026-08-28Current · 1 condition
?Unknown
No source reviewedPreview record
Report this result

Desktop

13 products

CLI

9 products

Unknown means no public evidence has been reviewed for that product and capability. It does not mean unsupported.

How statuses are assigned

Definition and scope

What this capability means

This row asks whether an uploaded or selected video contributes content to the model. Evidence should record whether the product uses sampled frames, native temporal input, extracted audio, a generated transcript, metadata, or some combination when that behavior is documented.

Record accepted containers and codecs, maximum bytes and duration, frame or sampling policy, audio handling, timecode awareness, resolution changes, model and plan restrictions, processing latency, and whether links are supported in addition to local uploads. A harness that accepts a video only for storage, sharing, or an unrelated editing tool remains unsupported for this row.

Traceable compatibility

Assertion ledger

Documentation evidence only. No runtime conformance test is implied.

Grokweb · current
Supported
Target
2026-08-28 Grok Web documentation observation · hosted-observation
Environment
hosted-default
Observed
2026-08-28
  • runtimetranscription and interpretation are documented, but frame sampling, temporal reasoning, timecodes, and audio-visual alignment are not
Evidence
Grok Botdesktop · current
Supported
Target
2026-08-28 Grok Bot desktop documentation observation · hosted-observation
Environment
hosted-default
Observed
2026-08-28
  • runtimevideo is accepted up to 200 MB, but frame, audio, transcript, timing, metadata, and sampling semantics are not documented on the reviewed page
Evidence
Geminiweb · current
Supported
Target
2026-08-28 Gemini Apps documentation observation · hosted-observation
Environment
hosted-default
Observed
2026-08-28
  • runtimedirect video upload and analysis are documented, with a 2 GB per-video limit
  • plantotal video is limited to 5 minutes without a Google AI plan and 1 hour with Google AI Pro or Ultra
  • runtimeframe sampling, temporal precision, timecodes, and audio-visual alignment are not established by the reviewed page
Evidence
Perplexityweb · current
Partial
Target
2026-08-28 Perplexity web documentation observation · hosted-observation
Environment
hosted-default
Observed
2026-08-28
  • runtimeMP4, MPEG, MOV, AVI, FLV, MPG, WebM, WMV, and 3GPP uploads are accepted and spoken content is automatically transcribed
  • runtimevisual scenes in video are explicitly not indexed or searchable, so this does not establish frame-based or temporal visual understanding
  • runtimethe general file-upload page documents a 40 MB per-file limit
Evidence
Gemini CLIcli · current
Supported
Target
current Gemini CLI custom-command documentation · dated-documentation
Environment
local-default
Observed
2026-08-28
  • runtimesupported video files referenced with @{...} inside a custom command are encoded and injected as multimodal input
  • runtimeaccepted containers, codecs, limits, frame sampling, timecodes, temporal reasoning, audio handling, and ordinary-prompt attachment methods are not established by the reviewed page
Evidence
  1. 1. Evidence checked 2026-08-28: xAI documents direct video uploads in Grok chats and describes transcription and interpretation of audio and video, but the reviewed FAQ does not specify frame sampling, temporal reasoning, or audio-visual alignment.
  2. 2. Evidence checked 2026-08-28: Grok Bot lists video among common supported inputs and documents a 200 MB per-video limit, but it does not describe which frames, audio, transcript, timing, or metadata reach the model.
  3. 3. Evidence checked 2026-08-28: Gemini Apps accept video uploads up to 2 GB and document total video-duration limits of 5 minutes without a Google AI plan or 1 hour with Google AI Pro or Ultra. The reviewed page does not define frame sampling, timecode, or audio-visual alignment semantics.
  4. 4. Evidence checked 2026-08-28: Perplexity web accepts common video containers and automatically transcribes spoken content, but its documentation explicitly says visual scenes in video are not indexed or searchable. This is transcript-oriented input rather than documented visual-video understanding.
  5. 5. Evidence checked 2026-08-28: Gemini CLI custom commands encode a supported video path referenced with @{...} and inject it as multimodal input. The reviewed page does not enumerate containers, codecs, duration, frame sampling, timecodes, or audio-visual alignment.