> ## Documentation Index
> Fetch the complete documentation index at: https://docs.laozhang.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Video understanding API: Gemini MP4 analysis

> Analyze complete MP4 videos, frames and audio together, with Gemini through LaoZhang API: uploads, URLs, chapters, JSON, clips, frame rate, and the OpenAI SDK.

Gemini models read a complete MP4 file, frames and audio track together. Use them to build chapters and summaries, transcribe narration, find what happens at a given moment, or compare videos. Upload the video as is; you do not need to extract frames first.

| Item          | Details                                                                                                              |
| ------------- | -------------------------------------------------------------------------------------------------------------------- |
| Endpoint      | `POST https://api.laozhang.ai/v1beta/models/{model}:generateContent` ([native Gemini API](/en/api-reference/gemini)) |
| Video source  | A local file inline with `inlineData`, or a public direct link or YouTube video with `fileData`                      |
| Billing       | Frames and audio both count as input tokens at the model's price; see the [model catalog](/en/models)                |
| Example model | `gemini-3.8-flash`                                                                                                   |

## Choose a model

Only Gemini models accept video; OpenAI models such as GPT-6 and GPT-5.6 do not. All of these models are usage-billed in the `default` group:

| Model ID                                   | Good for                                                       |
| ------------------------------------------ | -------------------------------------------------------------- |
| `gemini-3.8-flash`                         | The default choice: fast and suited to most video analysis     |
| `gemini-3.7-flash`                         | General video analysis, similar to 3.8 Flash                   |
| `gemini-3.1-pro-preview`, `gemini-2.5-pro` | Complex videos that need stronger reasoning, at a higher price |
| `gemini-3.5-flash-lite`                    | The lowest price, for bulk summaries and classification        |

## Set up your key and SDK

Create an API key in [Token management](https://api.laozhang.ai/token) with its billing mode set to Usage first (recommended) or Usage-based, and export it. The Python examples use the Google Gen AI SDK:

```bash theme={null}
export LAOZHANG_API_KEY="YOUR_LAOZHANG_API_KEY"
python -m pip install google-genai
```

LaoZhang API does not provide the Google File API (`/upload/v1beta/files`), so the SDK's `client.files.upload()` does not work. Send local videos inline with `inlineData` as shown below, and pass videos that already have a public URL with `fileData`.

## Quick start: send a local video

This example reads `input.mp4` from the current directory and asks the model to describe the scenes in order and transcribe the speech:

<Tabs>
  <Tab title="Python SDK">
    ```python theme={null}
    import os
    from pathlib import Path
    from google import genai
    from google.genai import types

    client = genai.Client(
        api_key=os.environ["LAOZHANG_API_KEY"],
        http_options={
            "base_url": "https://api.laozhang.ai",
            "api_version": "v1beta",
            "timeout": 300_000,
        },
    )
    video = types.Part.from_bytes(
        data=Path("input.mp4").read_bytes(),
        mime_type="video/mp4",
    )
    response = client.models.generate_content(
        model="gemini-3.8-flash",
        contents=[video, "Describe the video scene by scene and transcribe any speech."],
    )
    print(response.text)
    ```
  </Tab>

  <Tab title="cURL">
    A Base64 video is too long for the command line, so write the request body with the Python standard library first:

    ```python theme={null}
    import base64
    import json
    from pathlib import Path

    request = {
        "contents": [{
            "role": "user",
            "parts": [
                {"inlineData": {
                    "mimeType": "video/mp4",
                    "data": base64.b64encode(Path("input.mp4").read_bytes()).decode("ascii"),
                }},
                {"text": "Describe the video scene by scene and transcribe any speech."},
            ],
        }],
    }
    Path("video-request.json").write_text(json.dumps(request), encoding="utf-8")
    ```

    Then send it:

    ```bash theme={null}
    curl --fail-with-body --max-time 300 \
      "https://api.laozhang.ai/v1beta/models/gemini-3.8-flash:generateContent" \
      -H "Authorization: Bearer $LAOZHANG_API_KEY" \
      -H "Content-Type: application/json" \
      --data-binary @video-request.json
    ```
  </Tab>
</Tabs>

The answer is the `text` in `candidates[].content.parts`; the SDK exposes it as `response.text`.

When you send a video inline:

* Set `mimeType` to the file's real type, `video/mp4` for MP4. `data` holds the raw Base64 string with no `data:` prefix.
* Base64 makes the request about a third larger than the file, so larger videos take longer to upload. Allow a generous timeout.
* For large files or videos that are already hosted publicly, [pass a URL](#analyze-a-video-url) instead and skip the upload.

The examples below reuse the `client` created above, and the local video is always `input.mp4`.

## Example: split a video into chapters

State the time format and the fields you want for each chapter. The model uses both scene changes and speech to find the boundaries:

```python theme={null}
video = types.Part.from_bytes(
    data=Path("input.mp4").read_bytes(), mime_type="video/mp4"
)
response = client.models.generate_content(
    model="gemini-3.8-flash",
    contents=[
        video,
        "Split the video into chapters. For each chapter, give the start and end time (MM:SS), the title shown on screen, and a one-sentence summary of the narration.",
    ],
)
print(response.text)
```

The answer looks like this:

```text theme={null}
1. 00:00 - 00:06  Title: INTRO    Summary: Welcomes viewers to the product tour.
2. 00:06 - 00:12  Title: PRICING  Summary: The starter plan costs $29 per month.
3. 00:12 - 00:18  Title: DEMO     Summary: Shows how to upload a file and run the first job.
```

Frames and audio are aligned in time, so the model can tell you what was said while a given scene was on screen. Reported times can be about one second off from the actual cut; see [How accurate are the timestamps?](#how-accurate-are-the-timestamps) when you need exact seconds.

## Example: return structured JSON

To feed the result into a database or subtitle system, set `response_mime_type` and `response_schema` in `config`. The model returns JSON that matches the schema, and the SDK parses it into `response.parsed`:

```python theme={null}
from pydantic import BaseModel

class Chapter(BaseModel):
    start: str
    end: str
    title: str
    summary: str

video = types.Part.from_bytes(
    data=Path("input.mp4").read_bytes(), mime_type="video/mp4"
)
response = client.models.generate_content(
    model="gemini-3.8-flash",
    contents=[video, "List the chapters of this video. Use MM:SS for times."],
    config=types.GenerateContentConfig(
        response_mime_type="application/json",
        response_schema=list[Chapter],
    ),
)
for chapter in response.parsed:
    print(chapter.start, chapter.end, chapter.title, chapter.summary)
```

In a REST request, put the same settings in `generationConfig`:

```json theme={null}
"generationConfig": {
  "responseMimeType": "application/json",
  "responseSchema": {
    "type": "ARRAY",
    "items": {
      "type": "OBJECT",
      "properties": {
        "start": {"type": "STRING"},
        "end": {"type": "STRING"},
        "title": {"type": "STRING"},
        "summary": {"type": "STRING"}
      },
      "required": ["start", "end", "title", "summary"]
    }
  }
}
```

## Example: analyze one segment

Set `startOffset` and `endOffset` in `videoMetadata` to clip the video. The model processes only that segment, and usage is based on its length, which is cheaper than sending a long video in full when you care about a few minutes:

```python theme={null}
clip = types.Part(
    inline_data=types.Blob(
        data=Path("input.mp4").read_bytes(), mime_type="video/mp4"
    ),
    video_metadata=types.VideoMetadata(start_offset="12s", end_offset="18s"),
)
response = client.models.generate_content(
    model="gemini-3.8-flash",
    contents=[clip, "What is shown on screen in this segment, and what is said?"],
)
print(response.text)
```

In a REST request, `videoMetadata` goes in the same part as `inlineData` or `fileData`:

```json theme={null}
{
  "inlineData": {"mimeType": "video/mp4", "data": "<BASE64_DATA>"},
  "videoMetadata": {"startOffset": "12s", "endOffset": "18s"}
}
```

To ask about a single moment, you can also skip clipping and name the time in the prompt, for example "What is on screen at 00:14?"

## Analyze a video URL

If the video already has a public URL, pass it in `fileData` and the service downloads it, so you skip the Base64 upload. This example analyzes the first 60 seconds of a public YouTube video:

<Tabs>
  <Tab title="Python SDK">
    ```python theme={null}
    video = types.Part(
        file_data=types.FileData(
            file_uri="https://www.youtube.com/watch?v=9hE5-98ZeCg",
            mime_type="video/mp4",
        ),
        video_metadata=types.VideoMetadata(start_offset="0s", end_offset="60s"),
    )
    response = client.models.generate_content(
        model="gemini-3.8-flash",
        contents=[video, "Summarize what is shown and said in this clip, in order."],
    )
    print(response.text)
    ```
  </Tab>

  <Tab title="cURL">
    ```bash theme={null}
    curl --fail-with-body --max-time 300 \
      "https://api.laozhang.ai/v1beta/models/gemini-3.8-flash:generateContent" \
      -H "Authorization: Bearer $LAOZHANG_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "contents": [{
          "role": "user",
          "parts": [
            {
              "fileData": {
                "fileUri": "https://www.youtube.com/watch?v=9hE5-98ZeCg",
                "mimeType": "video/mp4"
              },
              "videoMetadata": {"startOffset": "0s", "endOffset": "60s"}
            },
            {"text": "Summarize what is shown and said in this clip, in order."}
          ]
        }]
      }'
    ```
  </Tab>
</Tabs>

URL requirements:

* A public YouTube video or a direct MP4 link.
* The URL must download without a login, cookies, or a signature. If the service cannot fetch it, the request fails with a 400 that says `Cannot fetch content from the provided URL`.
* For a video in private storage, create a temporary public link, or download it and send it with `inlineData`.

## Example: compare several videos

Put one video part per video in the same message, and refer to them by position in the prompt:

```python theme={null}
first = types.Part.from_bytes(data=Path("a.mp4").read_bytes(), mime_type="video/mp4")
second = types.Part.from_bytes(data=Path("b.mp4").read_bytes(), mime_type="video/mp4")
response = client.models.generate_content(
    model="gemini-3.8-flash",
    contents=[first, second, "What does each video cover? List the main differences between the first and the second."],
)
print(response.text)
```

Each video is counted separately; total usage is the sum of all videos.

## Control the frame rate and usage

By default the model samples one frame per second and processes the full audio track. Change the sampling rate with `videoMetadata.fps`:

* For slow-moving video such as lectures or screen recordings, a lower value such as `0.5` roughly halves the frame tokens.
* For fast action or brief on-screen moments, a higher value such as `2` roughly doubles them.
* Audio tokens are not affected by `fps`.

```python theme={null}
video = types.Part(
    inline_data=types.Blob(
        data=Path("input.mp4").read_bytes(), mime_type="video/mp4"
    ),
    video_metadata=types.VideoMetadata(fps=0.5),
)
```

`fps` can share one `videoMetadata` with `startOffset` and `endOffset`.

## Stream the answer

For long answers, show text as it is generated. The SDK uses `generate_content_stream`; in REST, change the method to `streamGenerateContent` and add `?alt=sse`:

```python theme={null}
video = types.Part.from_bytes(
    data=Path("input.mp4").read_bytes(), mime_type="video/mp4"
)
for chunk in client.models.generate_content_stream(
    model="gemini-3.8-flash",
    contents=[video, "Summarize this video in order."],
):
    print(chunk.text or "", end="", flush=True)
print()
```

The SDK may print an `is not a valid FinishReason` warning while streaming. It does not affect the output.

## Call it with the OpenAI SDK

If your application is built on the OpenAI SDK, send the video in a `video_url` part in Chat Completions. Only Base64 data URLs are accepted:

```python theme={null}
import base64
import os
from pathlib import Path
from openai import OpenAI

video_data = base64.b64encode(Path("input.mp4").read_bytes()).decode("ascii")

client = OpenAI(
    api_key=os.environ["LAOZHANG_API_KEY"],
    base_url="https://api.laozhang.ai/v1",
    timeout=300.0,
    max_retries=0,
)
response = client.chat.completions.create(
    model="gemini-3.8-flash",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "Describe the video scene by scene and transcribe any speech."},
            {"type": "video_url", "video_url": {"url": f"data:video/mp4;base64,{video_data}"}},
        ],
    }],
)
print(response.choices[0].message.content)
```

Limits of this approach:

* Install it with `pip install openai`, and use a Gemini model for `model`.
* A public URL in `video_url`, or a video in a `file` part, is dropped, and the model may answer without having seen the video.
* Video URLs, clipping, and `fps` are not available; use the native API when you need them.
* `video_tokens` greater than 0 in the response `usage` confirms the video was read.

## Read usage and billing

In the native API, `usageMetadata.promptTokensDetails` reports frame and audio tokens separately:

```json theme={null}
"promptTokensDetails": [
  {"modality": "VIDEO", "tokenCount": 1980},
  {"modality": "AUDIO", "tokenCount": 750},
  {"modality": "TEXT", "tokenCount": 30}
]
```

Both are billed as input tokens. To use fewer tokens, clip the video, lower `fps`, or choose a lower-priced model. The actual charge for each request is in [call logs](/en/faq/call-logs).

## Troubleshooting

* **400 with "Cannot fetch content from the provided URL":** the URL cannot be downloaded directly. Use a public direct link, or send the file with `inlineData`.
* **A file upload returns a web page or 404:** the File API is not available. Use `inlineData` or a `fileData` URL.
* **200, but the answer has nothing to do with the video:** check that usage includes `VIDEO`. In the OpenAI format, make sure you used `video_url` with a data URL.
* **Timeout:** raise the client timeout for large videos, or clip the video or switch to a URL. Check [call logs](/en/faq/call-logs) before you retry.
* **400 or 500 saying `video_url` is invalid:** use a Gemini model; OpenAI models do not accept video.

## FAQ

### Can I send a complete MP4 file?

Yes. Base64-encode the whole MP4 into `inlineData`, and the model reads the frames in order along with the audio. No frame extraction is needed. If the video already has a public URL, `fileData` saves the upload.

### Which models understand video?

These models:

* `gemini-3.8-flash`, `gemini-3.7-flash`
* `gemini-3.5-flash-lite`
* `gemini-3.1-pro-preview`, `gemini-2.5-pro`

Before using another Gemini model, send a short clip and check that usage includes `VIDEO`.

### Does the model hear the audio?

Yes. The audio track is processed together with the frames, so the model can transcribe dialogue, summarize narration, and answer questions such as what was said while a scene was on screen. `AUDIO` in the usage details is the audio track's token count.

### How accurate are the timestamps?

Chapter order and content are reliable, but every model can place a start or end time about one second away from the actual cut. When you need exact seconds:

* Clip the relevant segment with `startOffset` and `endOffset` and ask again.
* Raise `fps` so the model sees more frames.
* Generate timestamped subtitles with `whisper-1` and check against the speech timing; see [Create subtitles and timestamps](/en/api-capabilities/audio-transcription#create-subtitles-and-timestamps).

### Is there a size or length limit?

Inline video makes the request about a third larger than the file, and larger requests upload more slowly and are more likely to time out. Pass large files as a public URL with `fileData`, and split long videos with `startOffset` and `endOffset`.

### Is the Google File API supported?

No. None of these are available:

* The `/upload/v1beta/files` upload endpoint
* `files/...` references
* The SDK's `client.files.upload()`

Send local videos with `inlineData` and public videos with `fileData`.

### Can I generate videos?

This page covers understanding existing videos. To generate video, use [Wan 2.7](/en/api-capabilities/wan-video-generation) or [Seedance 2.0](/en/api-capabilities/seedance2-video-generation).

## Related documentation

* [Gemini protocol](/en/api-reference/gemini)
* [Image understanding API](/en/api-capabilities/vision-understanding)
* [Audio transcription and speech](/en/api-capabilities/audio-transcription)
* [Google video understanding guide](https://ai.google.dev/gemini-api/docs/video-understanding)
* [Model catalog and pricing](/en/models)
* [Call logs](/en/faq/call-logs)
