Skip to main content
Cline connects to LaoZhang API through OpenAI Compatible: set the Base URL to https://api.laozhang.ai/v1, enter your key, and type a full model ID. Cline calls tools through Chat Completions, so pick a model from the list below rather than GPT-6.

Set it up

1

Install Cline

Search for Cline in the VS Code extension marketplace and install it. A Cline icon appears in the activity bar.
2

Open settings

Click the gear icon at the top right of the Cline panel.
3

Enter the provider details

  • API Provider: OpenAI Compatible
  • Base URL: https://api.laozhang.ai/v1
  • API Key: your LaoZhang API key
  • Model: the full model ID, such as deepseek-v4-pro
Don’t choose the OpenAI provider; it always connects to OpenAI’s own endpoint.
4

Adjust model settings (optional)

In the model configuration, set Context Window and Max Output Tokens for the model you chose, and check Image Support if it accepts images. If you’re unsure, keep the defaults.
5

Test with a small task

Ask Cline to create a file and write one line to it. If it does, chat and tool calls both work.

Choose a model

Cline needs a model that returns tool calls reliably. These models have been confirmed to call tools in streaming Chat Completions requests: Prices are in the model catalog.

FAQ

With GPT-6 I get “Function tools with reasoning_effort are not supported”

GPT-6 models can only call tools in Chat Completions with reasoning effort set to none; otherwise they need the Responses API. Cline’s OpenAI Compatible provider uses Chat Completions, so switch to a model from the table above. To code with GPT-6, use Codex CLI or OpenCode.

I get “Unsupported parameter: ‘max_tokens’”

The model doesn’t accept max_tokens, which is the case for GPT-6 models. Switch to a model from the table above.

One task used a lot of tokens

Each step sends the conversation history, open files, and tool definitions. Give the token a quota and break large tasks into smaller ones. Models with a cache-read price, such as deepseek-v4-pro, bill repeated context at the cache rate; see prompt cache billing.