What the web app adds
Common differences
- It doesn’t know recent news. The model’s knowledge stops at its training cutoff; the web app fills the gap with web search. With the API, call a search service yourself and put the results in context.
- It gets arithmetic or word counts wrong. The web app quietly runs code for calculations, while the API model works it out in its head. Give it a code-execution tool, or ask it to show its working.
- Answers are shorter or less structured. The web app’s system prompt sets structure and length. Put the style you want in your own system prompt.
- It forgets what you said earlier. Each API request is a new conversation. Send earlier turns in
messages, and use prompt caching to cut the cost of the repeated prefix. - Answers vary each time. That’s sampling randomness. Lower
temperatureor constrain the output format in the prompt. - Reasoning feels shallower. Web apps often default to higher reasoning effort. Raise
reasoning_effortor the equivalent and leave enough output budget; see max_tokens.
The model names the wrong model or version
Asking a model “Which model are you?” through the API often gets the version wrong. That doesn’t mean you called the wrong model:- A model is named after training finishes, so it never learned its own name; its training data only contains earlier model names, so it guesses one of those.
- The web app gets it right because its hidden system prompt tells the model who it is.
- The
modelparameter in your request selects the model; the model itself can’t read that field.
Confirm which model you called
Don’t rely on the model’s own answer. Check:- the
modelfield in the response JSON; - the model name and charge for the request in call logs.
Get closer to the web app with the API
1
Write a system prompt
Set the role, tone, output format, length, and limits. This step makes the biggest difference.
2
Keep the conversation history
Append each user message and model reply to
messages. As the conversation grows, summarize it or keep only recent turns plus key facts.3
Add tools you need
Add search for current information, code execution for exact calculations, and retrieval for internal documents. Tool-call formats are covered in Chat Completions and Claude protocol.
4
Set parameters explicitly
Don’t rely on defaults; set
temperature, the output limit, and reasoning effort.https://api.laozhang.ai/v1 and enter your LaoZhang API key; see Invalid API key or 404 for setup tips.
The API can’t reproduce a web app exactly: vendors don’t publish their system prompts, some web features have no API, and web apps keep running experiments and switching models. In return, the API puts the prompt, settings, and context in your hands, so results are reproducible, which is what you need when building a product.