Skip to main content
Gemini models read a complete MP4 file, frames and audio track together. Use them to build chapters and summaries, transcribe narration, find what happens at a given moment, or compare videos. Upload the video as is; you do not need to extract frames first.

Choose a model

Only Gemini models accept video; OpenAI models such as GPT-6 and GPT-5.6 do not. All of these models are usage-billed in the default group:

Set up your key and SDK

Create an API key in Token management with its billing mode set to Usage first (recommended) or Usage-based, and export it. The Python examples use the Google Gen AI SDK:
LaoZhang API does not provide the Google File API (/upload/v1beta/files), so the SDK’s client.files.upload() does not work. Send local videos inline with inlineData as shown below, and pass videos that already have a public URL with fileData.

Quick start: send a local video

This example reads input.mp4 from the current directory and asks the model to describe the scenes in order and transcribe the speech:
The answer is the text in candidates[].content.parts; the SDK exposes it as response.text. When you send a video inline:
  • Set mimeType to the file’s real type, video/mp4 for MP4. data holds the raw Base64 string with no data: prefix.
  • Base64 makes the request about a third larger than the file, so larger videos take longer to upload. Allow a generous timeout.
  • For large files or videos that are already hosted publicly, pass a URL instead and skip the upload.
The examples below reuse the client created above, and the local video is always input.mp4.

Example: split a video into chapters

State the time format and the fields you want for each chapter. The model uses both scene changes and speech to find the boundaries:
The answer looks like this:
Frames and audio are aligned in time, so the model can tell you what was said while a given scene was on screen. Reported times can be about one second off from the actual cut; see How accurate are the timestamps? when you need exact seconds.

Example: return structured JSON

To feed the result into a database or subtitle system, set response_mime_type and response_schema in config. The model returns JSON that matches the schema, and the SDK parses it into response.parsed:
In a REST request, put the same settings in generationConfig:

Example: analyze one segment

Set startOffset and endOffset in videoMetadata to clip the video. The model processes only that segment, and usage is based on its length, which is cheaper than sending a long video in full when you care about a few minutes:
In a REST request, videoMetadata goes in the same part as inlineData or fileData:
To ask about a single moment, you can also skip clipping and name the time in the prompt, for example “What is on screen at 00:14?”

Analyze a video URL

If the video already has a public URL, pass it in fileData and the service downloads it, so you skip the Base64 upload. This example analyzes the first 60 seconds of a public YouTube video:
URL requirements:
  • A public YouTube video or a direct MP4 link.
  • The URL must download without a login, cookies, or a signature. If the service cannot fetch it, the request fails with a 400 that says Cannot fetch content from the provided URL.
  • For a video in private storage, create a temporary public link, or download it and send it with inlineData.

Example: compare several videos

Put one video part per video in the same message, and refer to them by position in the prompt:
Each video is counted separately; total usage is the sum of all videos.

Control the frame rate and usage

By default the model samples one frame per second and processes the full audio track. Change the sampling rate with videoMetadata.fps:
  • For slow-moving video such as lectures or screen recordings, a lower value such as 0.5 roughly halves the frame tokens.
  • For fast action or brief on-screen moments, a higher value such as 2 roughly doubles them.
  • Audio tokens are not affected by fps.
fps can share one videoMetadata with startOffset and endOffset.

Stream the answer

For long answers, show text as it is generated. The SDK uses generate_content_stream; in REST, change the method to streamGenerateContent and add ?alt=sse:
The SDK may print an is not a valid FinishReason warning while streaming. It does not affect the output.

Call it with the OpenAI SDK

If your application is built on the OpenAI SDK, send the video in a video_url part in Chat Completions. Only Base64 data URLs are accepted:
Limits of this approach:
  • Install it with pip install openai, and use a Gemini model for model.
  • A public URL in video_url, or a video in a file part, is dropped, and the model may answer without having seen the video.
  • Video URLs, clipping, and fps are not available; use the native API when you need them.
  • video_tokens greater than 0 in the response usage confirms the video was read.

Read usage and billing

In the native API, usageMetadata.promptTokensDetails reports frame and audio tokens separately:
Both are billed as input tokens. To use fewer tokens, clip the video, lower fps, or choose a lower-priced model. The actual charge for each request is in call logs.

Troubleshooting

  • 400 with “Cannot fetch content from the provided URL”: the URL cannot be downloaded directly. Use a public direct link, or send the file with inlineData.
  • A file upload returns a web page or 404: the File API is not available. Use inlineData or a fileData URL.
  • 200, but the answer has nothing to do with the video: check that usage includes VIDEO. In the OpenAI format, make sure you used video_url with a data URL.
  • Timeout: raise the client timeout for large videos, or clip the video or switch to a URL. Check call logs before you retry.
  • 400 or 500 saying video_url is invalid: use a Gemini model; OpenAI models do not accept video.

FAQ

Can I send a complete MP4 file?

Yes. Base64-encode the whole MP4 into inlineData, and the model reads the frames in order along with the audio. No frame extraction is needed. If the video already has a public URL, fileData saves the upload.

Which models understand video?

These models:
  • gemini-3.8-flash, gemini-3.7-flash
  • gemini-3.5-flash-lite
  • gemini-3.1-pro-preview, gemini-2.5-pro
Before using another Gemini model, send a short clip and check that usage includes VIDEO.

Does the model hear the audio?

Yes. The audio track is processed together with the frames, so the model can transcribe dialogue, summarize narration, and answer questions such as what was said while a scene was on screen. AUDIO in the usage details is the audio track’s token count.

How accurate are the timestamps?

Chapter order and content are reliable, but every model can place a start or end time about one second away from the actual cut. When you need exact seconds:
  • Clip the relevant segment with startOffset and endOffset and ask again.
  • Raise fps so the model sees more frames.
  • Generate timestamped subtitles with whisper-1 and check against the speech timing; see Create subtitles and timestamps.

Is there a size or length limit?

Inline video makes the request about a third larger than the file, and larger requests upload more slowly and are more likely to time out. Pass large files as a public URL with fileData, and split long videos with startOffset and endOffset.

Is the Google File API supported?

No. None of these are available:
  • The /upload/v1beta/files upload endpoint
  • files/... references
  • The SDK’s client.files.upload()
Send local videos with inlineData and public videos with fileData.

Can I generate videos?

This page covers understanding existing videos. To generate video, use Wan 2.7 or Seedance 2.0.