Skip to main content
Codex CLI connects to LaoZhang API through a custom provider in config.toml that points at the LaoZhang API host and uses the Responses API. You can then use GPT-6 and other OpenAI models to edit files and run commands, billed by token.

Install Codex CLI

Install it globally with npm (Node.js required):

Configure LaoZhang API

1

Set the key as an environment variable

Add the key to your shell profile so every new terminal can read it:
Use ~/.bashrc for bash. In Windows PowerShell, run setx LAOZHANG_API_KEY "YOUR_LAOZHANG_API_KEY" and open a new terminal.
2

Edit config.toml

Open or create ~/.codex/config.toml and add:
  • model_provider must match the name in [model_providers.laozhang].
  • Don’t name the provider openai, ollama, or lmstudio; Codex reserves those.
  • env_key takes the name of the environment variable, not the key itself.
3

Send a test request

From any directory, run:
If it prints connected, the setup works. Run codex in a project directory to start an interactive session.

Choose a model

Codex CLI calls models through the Responses API, so GPT-6 models can call tools without any reasoning-effort change. To switch for one session, add -m, for example codex -m gpt-6-astra. For model IDs Codex doesn’t recognize, such as gpt-5.4-mini, it warns Model metadata ... not found. The model still works, but Codex falls back to default values for details like context length. Prices are in the model catalog.

FAQ

I get a 401 or a missing environment variable error

Codex reads the key only from the variable named in env_key. In the same terminal you run Codex from, run test -n "$LAOZHANG_API_KEY" && echo set to confirm the variable is set. After editing your shell profile, open a new terminal.

I get a 404 or “stream disconnected”

Make sure base_url is https://api.laozhang.ai/v1, without a trailing / or /responses. If it still fails, check call logs to see whether the request arrived, and see Invalid API key or 404.

One task used a lot of tokens. Is that normal?

Yes. Each step sends the conversation history, file contents, and tool definitions, so a task can use hundreds of thousands of tokens. Give the token a quota and check each request’s usage in call logs. Repeated context is billed at the cache rate; see prompt cache billing.