The catalog is the source of truth for which models have a cache price and what it is. You don’t need to turn caching on, and there’s no separate cache-write fee.
Raise your hit rate
Caching matches on prefixes: only the part of a request that starts exactly like an earlier one can hit. When you build requests:- Put content that doesn’t change first: system prompts, tool definitions, long documents.
- Put content that changes every time last: the user’s question, timestamps, random IDs.
- In multi-turn chats, append new messages at the end and leave earlier history untouched.
- Keep one model per workload. Caches are separate per model and aren’t shared when you switch.
prompt_cache_key to group similar requests so they hit more often; see OpenAI prompt caching.
Confirm a cache hit
Check the cache field in the response’susage; any value above 0 is a hit:
Cached tokens are billed at the cache-read price and the rest of the input at the normal input price; the full formula is in Read the charges in a log entry. Call logs show the actual charge.
Things to keep in mind
- Gemini’s implicit cache hits unpredictably, because the upstream provider decides when it hits. Budget at uncached prices and treat hits as a bonus.
- Grok doesn’t guarantee hits upstream either, so budget the same way.
- Short prefixes don’t trigger caching. OpenAI’s minimum is 1,024 tokens; other vendors document their own thresholds.
- With long contexts, a cache hit can make the charge much lower than a full-price estimate from
prompt_tokens. That’s expected.