max_tokens caps how many tokens a model can generate in one reply. LaoZhang API passes it to the upstream model unchanged and adds no limit of its own; if you leave it out, the model’s default applies. Each protocol names the parameter differently, and the wrong name either fails or is ignored.
How the value affects a request
- Too low: the answer is cut off partway.
- Too high: the model doesn’t write more because of it and still stops when it’s done. A higher cap does raise the amount reserved before the request runs, which can get a request rejected when your balance is low.
- Not set: the model’s default applies. Defaults differ by vendor; some allow output until the context window is full, others default to a small cap.
Set it in each protocol
- Chat Completions
- Responses
- Anthropic Messages
- Gemini native
OpenAI has deprecated If another compatible model rejects
max_tokens in Chat Completions. Reasoning models such as GPT-6, GPT-5.x, and the o series need max_completion_tokens:max_completion_tokens, use max_tokens; the model’s guide and the error message tell you which one it expects.client for its protocol; the base URLs are in Connect your application to LaoZhang API.
Suggested values
Set the limit explicitly on every request so different model defaults don’t make output length unpredictable. Common starting points:- chat: 2,048–4,096
- code generation: 4,096–8,192
- long-form writing: 8,192–16,384, or stream the output
How it affects pre-authorized charges
Before a request runs, LaoZhang API reserves an estimated charge based on the model’s price, the input size, and the expected output, then settles on actual usage. An explicitmax_tokens lowers the expected output and therefore the reserved amount. With a long input and a small balance, a sensible output cap keeps the request from being rejected before it starts. See Why a request can fail with remaining balance.
FAQ
Does LaoZhang API limit or rewrite max_tokens?
No. The value goes to the upstream model unchanged; the only limit is the model’s own maximum output.My output was cut off. What should I do?
Check the signal in the table above to confirm the cap was hit. If it was:- Raise the cap.
- For reasoning models, lower the reasoning effort to leave room for the answer.
- Make sure the parameter name matches the protocol; reasoning models in Chat Completions need
max_completion_tokens.