> ## Documentation Index
> Fetch the complete documentation index at: https://docs.laozhang.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Does LaoZhang API bill prompt caching?

> Which LaoZhang API models have cache-read prices, how to raise your cache hit rate, and how to confirm hits from the usage fields.

Yes. Caching parameters go to the upstream provider unchanged, cache-hit fields come back in the response unchanged, and cached input is billed at the cache-read price in the [model catalog](/en/models). Models whose cache-read column shows "—" have no cache price, so all input is billed at the normal input price.

| Vendor | How caching triggers | Models with a cache-read price |
| - | - | - |
| OpenAI | Automatic, prefixes of 1,024 tokens or more | Most models, including GPT-6 and GPT-5.x |
| Google Gemini | Implicit caching, on by default | Gemini 3.x Flash and others; hit rate varies |
| DeepSeek | Automatic prefix matching | `deepseek-v4-flash`, `deepseek-v4-pro`, and others |
| xAI Grok | Automatic prefix matching | Some models; hits aren't guaranteed upstream |
| Moonshot Kimi | Automatic | `kimi-k2.6`, `kimi-k2.5` |

The catalog is the source of truth for which models have a cache price and what it is. You don't need to turn caching on, and there's no separate cache-write fee.

## Raise your hit rate

Caching matches on prefixes: only the part of a request that starts exactly like an earlier one can hit. When you build requests:

* Put content that doesn't change first: system prompts, tool definitions, long documents.
* Put content that changes every time last: the user's question, timestamps, random IDs.
* In multi-turn chats, append new messages at the end and leave earlier history untouched.
* Keep one model per workload. Caches are separate per model and aren't shared when you switch.

OpenAI models also accept `prompt_cache_key` to group similar requests so they hit more often; see [OpenAI prompt caching](https://platform.openai.com/docs/guides/prompt-caching).

## Confirm a cache hit

Check the cache field in the response's `usage`; any value above 0 is a hit:

| Endpoint | Hit field |
| - | - |
| Chat Completions | `usage.prompt_tokens_details.cached_tokens` |
| Responses | `usage.input_tokens_details.cached_tokens` |
| Gemini native | `usageMetadata.cachedContentTokenCount` |
| DeepSeek | `usage.prompt_cache_hit_tokens` |

Cached tokens are billed at the cache-read price and the rest of the input at the normal input price; the full formula is in [Read the charges in a log entry](/en/faq/call-logs#read-the-charges-in-a-log-entry). Call logs show the actual charge.

## Things to keep in mind

* **Gemini's implicit cache hits unpredictably**, because the upstream provider decides when it hits. Budget at uncached prices and treat hits as a bonus.
* **Grok doesn't guarantee hits upstream** either, so budget the same way.
* Short prefixes don't trigger caching. OpenAI's minimum is 1,024 tokens; other vendors document their own thresholds.
* With long contexts, a cache hit can make the charge much lower than a full-price estimate from `prompt_tokens`. That's expected.

## FAQ

### Do I need to enable caching in my request?

Not for the vendors above; their caching is automatic. LaoZhang API doesn't rewrite these requests or charge a cache-write fee.

### Why did an identical request miss the cache the second time?

Upstream caches expire after a period without use and may be cleared early at peak times, and a single changed character in the prefix causes a miss. Check whether the start of the request includes a timestamp, random content, or tool definitions in a changing order.

## Related pages

* [Model catalog and pricing](/en/models)
* [Read the charges in a log entry](/en/faq/call-logs#read-the-charges-in-a-log-entry)
* [Billing, token groups, and invoices](/en/pricing)
* [Google: context caching](https://ai.google.dev/gemini-api/docs/caching)
* [DeepSeek: context caching](https://api-docs.deepseek.com/guides/kv_cache)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.