> ## Documentation Index
> Fetch the complete documentation index at: https://docs.laozhang.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Why does the API differ from the web app?

> Why a model seems weaker through LaoZhang API or names the wrong version, how to confirm which model you called, and how to get closer to the web app.

It's the same model. Official web apps wrap it in a hidden system prompt, web search, memory, and tuned settings, while the API gives you only the model. LaoZhang API forwards requests unchanged, without injecting prompts, so it behaves like the official API.

## What the web app adds

| Capability | Official web app | Calling the API |
| - | - | - |
| System prompt | Injected every turn, not published | None until you write one |
| Web search | Triggered automatically | Off by default |
| Math and code execution | Built-in sandbox | You add a tool |
| Memory | History saved for you | Stateless; send prior turns each time |
| Long conversations | Summarized and trimmed for you | You truncate or summarize |
| Settings and reasoning effort | Tuned for you; some apps switch models automatically | Defaults apply |
| Formatting | Rendered Markdown, citations, and code highlighting | Plain text or JSON |

## Common differences

* **It doesn't know recent news.** The model's knowledge stops at its training cutoff; the web app fills the gap with web search. With the API, call a search service yourself and put the results in context.
* **It gets arithmetic or word counts wrong.** The web app quietly runs code for calculations, while the API model works it out in its head. Give it a code-execution tool, or ask it to show its working.
* **Answers are shorter or less structured.** The web app's system prompt sets structure and length. Put the style you want in your own system prompt.
* **It forgets what you said earlier.** Each API request is a new conversation. Send earlier turns in `messages`, and use [prompt caching](/en/faq/prompt-cache-billing) to cut the cost of the repeated prefix.
* **Answers vary each time.** That's sampling randomness. Lower `temperature` or constrain the output format in the prompt.
* **Reasoning feels shallower.** Web apps often default to higher reasoning effort. Raise `reasoning_effort` or the equivalent and leave enough output budget; see [max\_tokens](/en/faq/max-tokens).

## The model names the wrong model or version

Asking a model "Which model are you?" through the API often gets the version wrong. That doesn't mean you called the wrong model:

* A model is named after training finishes, so it never learned its own name; its training data only contains earlier model names, so it guesses one of those.
* The web app gets it right because its hidden system prompt tells the model who it is.
* The `model` parameter in your request selects the model; the model itself can't read that field.

If you need it to identify itself correctly, say so in the system prompt:

```python theme={null}
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["LAOZHANG_API_KEY"],
    base_url="https://api.laozhang.ai/v1",
)

response = client.chat.completions.create(
    model="gpt-6-sol",
    messages=[
        {"role": "system", "content": "You are GPT-6 Sol, a model made by OpenAI."},
        {"role": "user", "content": "Which model are you?"},
    ],
)
print(response.choices[0].message.content)
```

## Confirm which model you called

Don't rely on the model's own answer. Check:

1. the `model` field in the response JSON;
2. the model name and charge for the request in [call logs](https://api.laozhang.ai/log).

## Get closer to the web app with the API

<Steps>
  <Step title="Write a system prompt">
    Set the role, tone, output format, length, and limits. This step makes the biggest difference.
  </Step>

  <Step title="Keep the conversation history">
    Append each user message and model reply to `messages`. As the conversation grows, summarize it or keep only recent turns plus key facts.
  </Step>

  <Step title="Add tools you need">
    Add search for current information, code execution for exact calculations, and retrieval for internal documents. Tool-call formats are covered in [Chat Completions](/en/api-reference/chat-completions) and [Claude protocol](/en/api-reference/claude).
  </Step>

  <Step title="Set parameters explicitly">
    Don't rely on defaults; set `temperature`, the output limit, and reasoning effort.
  </Step>
</Steps>

If you'd rather not build this yourself, use a third-party client that already has these features. Set its API URL to `https://api.laozhang.ai/v1` and enter your LaoZhang API key; see [Invalid API key or 404](/en/faq/invalid-api-key) for setup tips.

The API can't reproduce a web app exactly: vendors don't publish their system prompts, some web features have no API, and web apps keep running experiments and switching models. In return, the API puts the prompt, settings, and context in your hands, so results are reproducible, which is what you need when building a product.

## Related pages

* [Text generation API: protocols and models](/en/api-capabilities/text-generation)
* [What is max\_tokens and what if I omit it?](/en/faq/max-tokens)
* [Check model availability and access](/en/faq/model-availability)
* [How to review and use call logs](/en/faq/call-logs)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.