Skip to main content
Prefix caching is on for every open source model. No configuration, no cache-write surcharge. When a request shares a prefix with earlier traffic (system prompt, tool definitions, conversation history), those tokens skip prefill and bill at the cached rate. Cached input is 80% off on GLM-5.2 and GLM-5.3-Flash, and 90% off on Kimi K3. Other open source models cache automatically too; their cached tokens currently bill at the regular input rate.

Reading cache hits

Every response reports how much of the prompt was served from cache:
cached_tokens billed at the cached rate, the remainder of prompt_tokens at the input rate.

Getting hits

Prompt caching: a request matching the stable prefix is a cache hit and reuses those tokens; changing any token in the prefix is a cache miss
Matching is exact-prefix and block-aligned. To maximize hit rate:
  • Put stable content first: system prompt, then tool definitions, then history. Variable content (the user’s latest message, retrieved context) goes last.
  • Keep the prefix byte-identical between turns. A timestamp or request ID in the system prompt kills every hit after it.
  • Short prompts rarely hit. Caching operates on ~1k-token blocks, so a 300-token prompt has nothing to reuse.
Multi-turn agent loops get this for free: each turn re-sends the previous turns verbatim, so everything but the newest turn is a cache hit.

Session key

Caching is automatic and needs no key. A key answers the other half of the question: which worker serves the turn. A conversation’s prefix lives in the cache of the worker that prefilled it, so a follow-up that lands on a different worker re-prefills the whole thing at full price. Send one id per conversation and every turn routes to the worker that already holds its prefix. Two carriers, both OpenAI convention, both passed through by OpenRouter:
A client that already carries its conversation id in a header sends the same value there instead:
Picking keys:
  • One key per conversation or agent session. Generate it when the conversation starts and reuse it for every turn, including retries of the same turn.
  • Never a shared app-wide or canary value. One key across unrelated traffic funnels all of it onto a single worker. That worker fills up, sheds the overflow with a 429, and you spend the cache savings on retries.
  • Keys are hashed server-side and never stored raw. What we forward is derived from the hash, so it carries none of your value.
  • A key changes placement, never billing. Rates and the cached-input discount are the same with or without it.
Values are opaque to us: any non-empty string up to 512 bytes. A value that is empty, the wrong type, or longer than that counts as absent, so the request is served without a pin rather than rejected. When a request carries both carriers, the header is the one used.
Rolling out per model, DeepSeek V4 Flash (morph-dsv4flash) first. Sending the key to any other model is harmless: it is read, recorded, and ignored until that model’s rollout completes.
run_id on Agent Runs is the heavier version of the same idea. A run id schedules a whole tool-calling run as one unit: sticky placement, priority resume, and whole-run admission, which under load means whole runs pause rather than every run getting slow. A session key does placement and nothing else. Use a run id for an agent run on Kimi K3, a session key for any other multi-turn conversation. If a request carries both, run_id wins: you named the run yourself, and the session key is only our read of where its prefix lives.

Cache TTL

By default cached prefixes persist under LRU eviction, with no fixed expiry. To control retention per request, pass cache_ttl: Expiry is sliding, Anthropic-style: every cache hit refreshes the clock. Past the TTL the prefix stops hitting entirely (full recompute), and re-sending it caches it fresh. A prefix shared by multiple requests keeps the longest surviving TTL.
cache_ttl is rolling out now, GLM-5.2 first and Kimi K3 at launch. Requests that include it are accepted today; the field takes effect as each model’s rollout completes.
Fine print:
  • Invalid cache_ttl values are rejected with a 400. Only the five tiers above are accepted.
  • Expiry granularity is ~30 seconds: treat a TTL as “at least this long, expired within ~30s after.”
  • Omitting cache_ttl keeps the default behavior (LRU, no fixed expiry).
  • A session key and cache_ttl are independent: the key picks the worker, the TTL controls how long that worker keeps the prefix.

Pitfalls

Your prefix is changing between requests. Diff two consecutive prompts byte-for-byte; the first divergent token ends the cacheable prefix. Common culprits: timestamps, UUIDs, or shuffled tool order in the system prompt.
Editing an earlier message invalidates everything after it. Expected: caching is prefix-based, so append, don’t rewrite.

See Also