Reusing the KV-cache for stable prompt prefixes — the cheapest optimization most apps skip.
Skip to solutionKEEP THE
mediumAI Engineering
What is prompt caching and how does it reduce cost and latency?
599 views
01
Understand the problem
prompt-cachingcostlatencykv-cache
02
Attempt it yourself
Sketch your approach before reading the solution — that's what interviews test.
Nudge consolestandby
Stuck? Beam a request up — the console returns a conceptual nudge that guides your logic without spoiling the implementation.
03
Study the solution
Providers cache the internal computation (KV cache) for a prompt prefix they have seen recently, so repeated calls that share that prefix skip recomputing it — cached input tokens are billed at a large discount (up to ~90%) and time-to-first-token drops. To exploit it, structure prompts as stable-prefix-first (system p
Solution ready — 2 min read
Classified // press E to declassify
04
Read the code
Explicit cache breakpoints (Anthropic style)
await client.messages.create({
model, max_tokens: 800,
system: [
{ type: "text", text: BIG_STABLE_RUBRIC, // 8k tokens, never changes
cache_control: { type: "ephemeral" } },
],
tools: TOOLS_SORTED_DETERMINISTICALLY,
messages: [
...priorTurns, // growing cached prefix
{ role: "user", content: newQuestion }, // only this is fresh compute
],
});05
Join the discussion
Discussion (0)
Sign in to join the discussion.
No responses yet. Be the first to share what you think.
Transmission complete // awaiting log
KEEP THE
STREAK ALIVE.
Dossier 37 of 80 decoded in the AI Engineering track. One more won't hurt.