Cached input is 80% off on GLM-5.2 and GLM-5.3-Flash, and 90% off on Kimi K3. Other open source models cache automatically too; their cached tokens currently bill at the regular input rate.
Reading cache hits
Every response reports how much of the prompt was served from cache:cached_tokens billed at the cached rate, the remainder of prompt_tokens at the input rate.
Getting hits

- Put stable content first: system prompt, then tool definitions, then history. Variable content (the user’s latest message, retrieved context) goes last.
- Keep the prefix byte-identical between turns. A timestamp or request ID in the system prompt kills every hit after it.
- Short prompts rarely hit. Caching operates on ~1k-token blocks, so a 300-token prompt has nothing to reuse.
Session key
Caching is automatic and needs no key. A key answers the other half of the question: which worker serves the turn. A conversation’s prefix lives in the cache of the worker that prefilled it, so a follow-up that lands on a different worker re-prefills the whole thing at full price. Send one id per conversation and every turn routes to the worker that already holds its prefix. Two carriers, both OpenAI convention, both passed through by OpenRouter:- cURL
- Python
- TypeScript
- One key per conversation or agent session. Generate it when the conversation starts and reuse it for every turn, including retries of the same turn.
- Never a shared app-wide or canary value. One key across unrelated traffic funnels all of it onto a single worker. That worker fills up, sheds the overflow with a 429, and you spend the cache savings on retries.
- Keys are hashed server-side and never stored raw. What we forward is derived from the hash, so it carries none of your value.
- A key changes placement, never billing. Rates and the cached-input discount are the same with or without it.
Rolling out per model, DeepSeek V4 Flash (
morph-dsv4flash) first. Sending the key to any other model is harmless: it is read, recorded, and ignored until that model’s rollout completes.run_id on Agent Runs is the heavier version of the same idea. A run id schedules a whole tool-calling run as one unit: sticky placement, priority resume, and whole-run admission, which under load means whole runs pause rather than every run getting slow. A session key does placement and nothing else. Use a run id for an agent run on Kimi K3, a session key for any other multi-turn conversation. If a request carries both, run_id wins: you named the run yourself, and the session key is only our read of where its prefix lives.
Cache TTL
By default cached prefixes persist under LRU eviction, with no fixed expiry. To control retention per request, passcache_ttl:
Expiry is sliding, Anthropic-style: every cache hit refreshes the clock. Past the TTL the prefix stops hitting entirely (full recompute), and re-sending it caches it fresh. A prefix shared by multiple requests keeps the longest surviving TTL.
cache_ttl is rolling out now, GLM-5.2 first and Kimi K3 at launch. Requests that include it are accepted today; the field takes effect as each model’s rollout completes.- cURL
- Python
- TypeScript
- Invalid
cache_ttlvalues are rejected with a 400. Only the five tiers above are accepted. - Expiry granularity is ~30 seconds: treat a TTL as “at least this long, expired within ~30s after.”
- Omitting
cache_ttlkeeps the default behavior (LRU, no fixed expiry). - A session key and
cache_ttlare independent: the key picks the worker, the TTL controls how long that worker keeps the prefix.
Pitfalls
cached_tokens is 0 on every request
cached_tokens is 0 on every request
Your prefix is changing between requests. Diff two consecutive prompts byte-for-byte; the first divergent token ends the cacheable prefix. Common culprits: timestamps, UUIDs, or shuffled tool order in the system prompt.
Hits stop after an edit mid-conversation
Hits stop after an edit mid-conversation
Editing an earlier message invalidates everything after it. Expected: caching is prefix-based, so append, don’t rewrite.
See Also
- Standby Requests — stack a 50% tier discount on top: cached standby input is $0.11/1M
- Agent Runs —
run_id, the whole-run version of a session key - Open Source Models — the models this page prices
- Compact — shrink context before caching it