Prompt caching lets an LLM provider reuse the computed state of a prompt prefix instead of reprocessing it, cutting the cost of repeated input tokens by roughly 90%. If you’re building agents, it’s the single highest-leverage line item on your bill, and the easiest one to silently break.
The realization clicked while instrumenting per-turn usage in aloud, my voice agent. Do the math on one modest session: a ~2K-token stable prefix (system prompt plus tool schemas) resent on every one of 30 turns is 60K input tokens before the conversation itself, which is also resent in full and grows every turn. Across a session the model generates maybe a tenth of what it bills. The rest is you paying full price to re-send bytes the provider has already processed.
Why is the agent loop quadratic?
LLM APIs are stateless. Turn N’s request contains all N-1 previous turns: the system prompt, every tool schema, every tool result. Sum that over a session and input cost is O(turns²) while output cost stays linear. This is why input tokens, not output tokens, dominate agent bills, and why caching exists at all: the provider keeps the attention keys and values (the KV cache) it computed for your prefix, and a matching prefix on the next request costs a lookup instead of a forward pass.
The one rule everything follows from
Caching is a prefix match. Any byte change anywhere in the prefix invalidates everything after it.
On Anthropic’s API the prompt renders in a fixed order (tools, then system, then messages), and the cache key is the exact bytes up to a breakpoint. One reordered JSON key, one datetime.now() interpolated into the system prompt, one tool added mid-session, and everything downstream is a miss. No error. The requests keep succeeding. The bill just doubles.
So all caching design reduces to one discipline: stable content first, volatile content last. Frozen system prompt and deterministic tool list at the front, per-session context next, the growing conversation after that, and anything per-request (timestamps, request IDs) at the very end or deleted.
How the three big providers implement it
| Anthropic | OpenAI | Gemini | |
|---|---|---|---|
| Mechanism | explicit cache_control breakpoints | automatic | implicit (auto) + explicit cache objects |
| Read discount | 90% | 75-90% | ~90% |
| Write premium | 1.25x (5-min TTL) / 2x (1-h) | none | none (implicit); storage $/hr (explicit) |
| Control | full | none | none / full |
Anthropic makes you place breakpoints (max 4 per request); a cache read refreshes the TTL for free, so requests less than five minutes apart keep the cache warm indefinitely. OpenAI and Gemini’s implicit mode hash your prefix automatically: no markers, no write premium, but also no control. The portable skill is the same everywhere. If your prompt assembly isn’t prefix-stable, no provider can help you.
My projects run Gemini (aloud on 2.5 Flash for voice, the football content agent on Vertex AI), which means implicit caching does the work, but only if I keep the prefix byte-stable, which is exactly the same discipline the explicit APIs teach.
The silent invalidators
Every one of these has burned someone. Grep your prompt-assembly path for:
datetime.now() in the system prompt → new prefix every request
uuid4() / request IDs early in content → same
json.dumps(...) without sort_keys=True → nondeterministic bytes
tools=build_tools(user) → per-user prefix, no sharing
if flag: system += "..." → a distinct prefix per flag combo
The fix is always the same: move it after the last breakpoint, make it deterministic, or delete it.
Verify it, then keep verifying it
The usage block on every response is the only ground truth. On Anthropic: cache_read_input_tokens (the win), cache_creation_input_tokens (the 1.25x write), and input_tokens (the uncached tail; total prompt size is the sum of all three). On Gemini: usageMetadata.cachedContentTokenCount.
A healthy agent loop reads everything accumulated so far and writes only what the last turn added. If writes are near the full conversation size every turn, something upstream of your breakpoint is rewriting bytes. The costliest caching failure in production isn’t a bad first implementation. It’s a regression six weeks later, when someone adds a dynamic field to the system prompt and the bill doubles with zero errors. My rule now: an integration test asserting a second identical request shows a nonzero cache read. Caching you don’t monitor is caching you don’t have.
What caching can’t fix
Two things still bite even with perfect prefix hygiene. Parallel fan-out: N simultaneous requests over the same fresh prefix all pay full price, because an entry only becomes readable after the first response starts streaming. Fire one, await the first token, then fire the rest. And history edits: anything that rewrites earlier messages (compaction, injected-then-deleted reminders) is by definition a prefix change. Newer Anthropic models have cache-preserving escape hatches, like mid-conversation system messages appended after the history instead of edits to the top-level system prompt, but the mental model stays: treat conversation history as append-only.
Beyond caching, the levers that actually move an agent bill, in the order I’d pull them: keep deterministic work (token counting, context assembly, retrieval) in plain code instead of LLM calls; run latency-insensitive pipelines through the Batch API (50% off, and it stacks with cache reads); and give background jobs to a small model while the hot path keeps the capable one, as separate calls, never as a serial router in front of the main model.
The one-line summary I’d give my past self: design the prompt like a git history, append-only and stable at the front, and read the usage numbers like you read your bank statement.
EOF · back to posts