Skip to main content
Yes. Caching parameters go to the upstream provider unchanged, cache-hit fields come back in the response unchanged, and cached input is billed at the cache-read price in the model catalog. Models whose cache-read column shows ”—” have no cache price, so all input is billed at the normal input price. The catalog is the source of truth for which models have a cache price and what it is. You don’t need to turn caching on, and there’s no separate cache-write fee.

Raise your hit rate

Caching matches on prefixes: only the part of a request that starts exactly like an earlier one can hit. When you build requests:
  • Put content that doesn’t change first: system prompts, tool definitions, long documents.
  • Put content that changes every time last: the user’s question, timestamps, random IDs.
  • In multi-turn chats, append new messages at the end and leave earlier history untouched.
  • Keep one model per workload. Caches are separate per model and aren’t shared when you switch.
OpenAI models also accept prompt_cache_key to group similar requests so they hit more often; see OpenAI prompt caching.

Confirm a cache hit

Check the cache field in the response’s usage; any value above 0 is a hit: Cached tokens are billed at the cache-read price and the rest of the input at the normal input price; the full formula is in Read the charges in a log entry. Call logs show the actual charge.

Things to keep in mind

  • Gemini’s implicit cache hits unpredictably, because the upstream provider decides when it hits. Budget at uncached prices and treat hits as a bonus.
  • Grok doesn’t guarantee hits upstream either, so budget the same way.
  • Short prefixes don’t trigger caching. OpenAI’s minimum is 1,024 tokens; other vendors document their own thresholds.
  • With long contexts, a cache hit can make the charge much lower than a full-price estimate from prompt_tokens. That’s expected.

FAQ

Do I need to enable caching in my request?

Not for the vendors above; their caching is automatic. LaoZhang API doesn’t rewrite these requests or charge a cache-write fee.

Why did an identical request miss the cache the second time?

Upstream caches expire after a period without use and may be cleared early at peak times, and a single changed character in the prefix causes a miss. Check whether the start of the request includes a timestamp, random content, or tool definitions in a changing order.