These settings and models were last verified on October 5, 2026, with a
default-group token in Usage-first mode. The model catalog and the console are the source of truth for availability and prices.
Base URL by protocol
For direct connections from Europe or North America, use
api-vip.laozhang.ai instead. If api.laozhang.ai can’t be reached, use api2.laozhang.ai. Paths and keys stay the same.
Claude models are temporarily offline due to limited capacity, and we’ll post an announcement when they return. Tools that only speak the Anthropic format, such as Claude Code, can currently call GLM-5.2 or DeepSeek V4 with a
default-group, Usage-first token; see Claude Code.Before you pick a model
Coding agents and clients with tools (MCP or function calling) ask the model to call tools, and models differ in what parameters they accept. The cases you’re most likely to hit:- GPT-6 models fail to call tools in Chat Completions with
Function tools with reasoning_effort are not supported. Tools that use the Responses API (Codex CLI, OpenCode) aren’t affected. For tools that use Chat Completions, such as Cline, pick a model from the list below or set reasoning effort tononein the tool. - GPT-6 models reject
max_tokenswithUnsupported parameter: 'max_tokens'. Turn off the client’s max-tokens setting or choose another model. - For tool calls through Chat Completions, use one of these models:
- OpenAI:
gpt-5.6-sol,gpt-5.4-mini - DeepSeek:
deepseek-v4-pro,deepseek-v4-flash - Qwen and GLM:
qwen3-coder-plus,glm-5.1 - Gemini:
gemini-3.1-pro-preview,gemini-3.8-flash
- OpenAI:
Steps for any tool
1
Create a token
Create a token in Tokens in the
default group with the Usage-first billing mode. Give tokens for coding agents a quota; see API key management.2
Copy the model ID
Copy the full model ID from the model catalog rather than typing it.
3
Enter the URL and key
Use the base URL for the tool’s protocol from the table above, and put the key in the tool’s settings or an environment variable.
4
Send a test message
Ask for a one-line reply first, then try a tool call such as writing a file. If something fails, read the error, then check call logs to see whether the request arrived.
https://api.laozhang.ai/v1. For invalid_api_key or 404 errors, see Invalid API key or 404.
FAQ
Coding agents use a lot of tokens. How do I keep costs down?
Each step sends the conversation history, file contents, and tool definitions, so one task can use hundreds of thousands of tokens.- Give a dedicated token a quota so it stops when the quota runs out.
- Use cheaper models such as
gpt-5.4-miniordeepseek-v4-flashfor simple tasks. - Choose models with a cache-read price so repeated context is billed at the cache rate; see prompt cache billing.
The tool says the model doesn’t exist or isn’t allowed
SendGET https://api.laozhang.ai/v1/models with the same token and confirm the model ID is in the list, then check the token’s group. See Check model availability and access.