> ## Documentation Index
> Fetch the complete documentation index at: https://docs.laozhang.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# What is max_tokens and what if I omit it?

> How the output limit works in Chat Completions, Responses, Anthropic Messages, and Gemini on LaoZhang API, and how to spot truncated output.

`max_tokens` caps how many tokens a model can generate in one reply. LaoZhang API passes it to the upstream model unchanged and adds no limit of its own; if you leave it out, the model's default applies. Each protocol names the parameter differently, and the wrong name either fails or is ignored.

| Protocol | Parameter | Required | Signal when the limit is hit |
| - | - | - | - |
| Chat Completions | `max_completion_tokens` (older `max_tokens`) | No | `finish_reason` is `length` |
| Responses | `max_output_tokens` | No | `status` is `incomplete`, reason `max_output_tokens` |
| Anthropic Messages | `max_tokens` | Yes | `stop_reason` is `max_tokens` |
| Gemini native | `generationConfig.maxOutputTokens` | No | `finishReason` is `MAX_TOKENS` |

## How the value affects a request

* **Too low**: the answer is cut off partway.
* **Too high**: the model doesn't write more because of it and still stops when it's done. A higher cap does raise the amount reserved before the request runs, which can get a request rejected when your balance is low.
* **Not set**: the model's default applies. Defaults differ by vendor; some allow output until the context window is full, others default to a small cap.

For reasoning models, the hidden reasoning also counts against the output budget even though it isn't in the reply. With a cap that's too low, the model can run out while still reasoning and return an empty or truncated answer.

## Set it in each protocol

<Tabs>
  <Tab title="Chat Completions">
    OpenAI has deprecated `max_tokens` in Chat Completions. Reasoning models such as GPT-6, GPT-5.x, and the o series need `max_completion_tokens`:

    ```python theme={null}
    response = client.chat.completions.create(
        model="gpt-6-sol",
        messages=[{"role": "user", "content": "Summarize this article"}],
        max_completion_tokens=4096,
    )
    print(response.choices[0].finish_reason)  # "length" means the cap was hit
    ```

    If another compatible model rejects `max_completion_tokens`, use `max_tokens`; the model's guide and the error message tell you which one it expects.
  </Tab>

  <Tab title="Responses">
    ```python theme={null}
    response = client.responses.create(
        model="gpt-6-sol",
        input="Summarize this article",
        max_output_tokens=4096,
    )
    print(response.status)  # if "incomplete", check response.incomplete_details
    ```
  </Tab>

  <Tab title="Anthropic Messages">
    `max_tokens` is required; leaving it out returns a 400:

    ```python theme={null}
    message = client.messages.create(
        model="glm-5.2",
        max_tokens=4096,
        messages=[{"role": "user", "content": "Summarize this article"}],
    )
    print(message.stop_reason)  # "max_tokens" means the cap was hit
    ```
  </Tab>

  <Tab title="Gemini native">
    ```json theme={null}
    {
      "contents": [{"parts": [{"text": "Summarize this article"}]}],
      "generationConfig": {"maxOutputTokens": 4096}
    }
    ```

    `candidates[0].finishReason` set to `MAX_TOKENS` means the cap was hit.
  </Tab>
</Tabs>

Create each `client` for its protocol; the base URLs are in [Connect your application to LaoZhang API](/en/api-manual).

## Suggested values

Set the limit explicitly on every request so different model defaults don't make output length unpredictable. Common starting points:

* chat: 2,048–4,096
* code generation: 4,096–8,192
* long-form writing: 8,192–16,384, or stream the output

Each model's maximum output is listed in its vendor's documentation. Above that maximum, the model either applies its own cap or returns a parameter error.

## How it affects pre-authorized charges

Before a request runs, LaoZhang API reserves an estimated charge based on the model's price, the input size, and the expected output, then settles on actual usage. An explicit `max_tokens` lowers the expected output and therefore the reserved amount. With a long input and a small balance, a sensible output cap keeps the request from being rejected before it starts. See [Why a request can fail with remaining balance](/en/faq/balance-insufficient).

## FAQ

### Does LaoZhang API limit or rewrite max\_tokens?

No. The value goes to the upstream model unchanged; the only limit is the model's own maximum output.

### My output was cut off. What should I do?

Check the signal in the table above to confirm the cap was hit. If it was:

1. Raise the cap.
2. For reasoning models, lower the reasoning effort to leave room for the answer.
3. Make sure the parameter name matches the protocol; reasoning models in Chat Completions need `max_completion_tokens`.

### Why is the reply short even with a high limit?

The limit is only a ceiling, and the model stops when it considers the answer complete. For longer output, state the length and structure you want in the prompt.

## Related pages

* [OpenAI protocol: Chat Completions](/en/api-reference/chat-completions)
* [Claude protocol: Anthropic Messages API](/en/api-reference/claude)
* [Gemini protocol: generateContent and SDK](/en/api-reference/gemini)
* [Why a request can fail with remaining balance](/en/faq/balance-insufficient)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.