Skip to main content
max_tokens caps how many tokens a model can generate in one reply. LaoZhang API passes it to the upstream model unchanged and adds no limit of its own; if you leave it out, the model’s default applies. Each protocol names the parameter differently, and the wrong name either fails or is ignored.

How the value affects a request

  • Too low: the answer is cut off partway.
  • Too high: the model doesn’t write more because of it and still stops when it’s done. A higher cap does raise the amount reserved before the request runs, which can get a request rejected when your balance is low.
  • Not set: the model’s default applies. Defaults differ by vendor; some allow output until the context window is full, others default to a small cap.
For reasoning models, the hidden reasoning also counts against the output budget even though it isn’t in the reply. With a cap that’s too low, the model can run out while still reasoning and return an empty or truncated answer.

Set it in each protocol

OpenAI has deprecated max_tokens in Chat Completions. Reasoning models such as GPT-6, GPT-5.x, and the o series need max_completion_tokens:
If another compatible model rejects max_completion_tokens, use max_tokens; the model’s guide and the error message tell you which one it expects.
Create each client for its protocol; the base URLs are in Connect your application to LaoZhang API.

Suggested values

Set the limit explicitly on every request so different model defaults don’t make output length unpredictable. Common starting points:
  • chat: 2,048–4,096
  • code generation: 4,096–8,192
  • long-form writing: 8,192–16,384, or stream the output
Each model’s maximum output is listed in its vendor’s documentation. Above that maximum, the model either applies its own cap or returns a parameter error.

How it affects pre-authorized charges

Before a request runs, LaoZhang API reserves an estimated charge based on the model’s price, the input size, and the expected output, then settles on actual usage. An explicit max_tokens lowers the expected output and therefore the reserved amount. With a long input and a small balance, a sensible output cap keeps the request from being rejected before it starts. See Why a request can fail with remaining balance.

FAQ

Does LaoZhang API limit or rewrite max_tokens?

No. The value goes to the upstream model unchanged; the only limit is the model’s own maximum output.

My output was cut off. What should I do?

Check the signal in the table above to confirm the cap was hit. If it was:
  1. Raise the cap.
  2. For reasoning models, lower the reasoning effort to leave room for the answer.
  3. Make sure the parameter name matches the protocol; reasoning models in Chat Completions need max_completion_tokens.

Why is the reply short even with a high limit?

The limit is only a ceiling, and the model stops when it considers the answer complete. For longer output, state the length and structure you want in the prompt.