Skip to main content
LaoZhang API speaks the OpenAI, Anthropic Messages, and Gemini protocols, so most AI tools only need a new base URL, your LaoZhang API key, and a model ID. The table below lists each tool’s protocol, recommended models, and setup guide. These settings and models were last verified on October 5, 2026, with a default-group token in Usage-first mode. The model catalog and the console are the source of truth for availability and prices.

Base URL by protocol

For direct connections from Europe or North America, use api-vip.laozhang.ai instead. If api.laozhang.ai can’t be reached, use api2.laozhang.ai. Paths and keys stay the same.
Claude models are temporarily offline due to limited capacity, and we’ll post an announcement when they return. Tools that only speak the Anthropic format, such as Claude Code, can currently call GLM-5.2 or DeepSeek V4 with a default-group, Usage-first token; see Claude Code.

Before you pick a model

Coding agents and clients with tools (MCP or function calling) ask the model to call tools, and models differ in what parameters they accept. The cases you’re most likely to hit:
  • GPT-6 models fail to call tools in Chat Completions with Function tools with reasoning_effort are not supported. Tools that use the Responses API (Codex CLI, OpenCode) aren’t affected. For tools that use Chat Completions, such as Cline, pick a model from the list below or set reasoning effort to none in the tool.
  • GPT-6 models reject max_tokens with Unsupported parameter: 'max_tokens'. Turn off the client’s max-tokens setting or choose another model.
  • For tool calls through Chat Completions, use one of these models:
    • OpenAI: gpt-5.6-sol, gpt-5.4-mini
    • DeepSeek: deepseek-v4-pro, deepseek-v4-flash
    • Qwen and GLM: qwen3-coder-plus, glm-5.1
    • Gemini: gemini-3.1-pro-preview, gemini-3.8-flash
Each of these models has been confirmed to return tool calls in a streaming request. For more on output limits, see max_tokens.

Steps for any tool

1

Create a token

Create a token in Tokens in the default group with the Usage-first billing mode. Give tokens for coding agents a quota; see API key management.
2

Copy the model ID

Copy the full model ID from the model catalog rather than typing it.
3

Enter the URL and key

Use the base URL for the tool’s protocol from the table above, and put the key in the tool’s settings or an environment variable.
4

Send a test message

Ask for a one-line reply first, then try a tool call such as writing a file. If something fails, read the error, then check call logs to see whether the request arrived.
If a tool has no custom-provider option, choose “OpenAI-compatible” or the OpenAI type and change the base URL to https://api.laozhang.ai/v1. For invalid_api_key or 404 errors, see Invalid API key or 404.

FAQ

Coding agents use a lot of tokens. How do I keep costs down?

Each step sends the conversation history, file contents, and tool definitions, so one task can use hundreds of thousands of tokens.
  • Give a dedicated token a quota so it stops when the quota runs out.
  • Use cheaper models such as gpt-5.4-mini or deepseek-v4-flash for simple tasks.
  • Choose models with a cache-read price so repeated context is billed at the cache rate; see prompt cache billing.

The tool says the model doesn’t exist or isn’t allowed

Send GET https://api.laozhang.ai/v1/models with the same token and confirm the model ID is in the list, then check the token’s group. See Check model availability and access.