Skip to solution
mediumAI Engineering

What is prompt caching and how does it reduce cost and latency?

599 views
01

Understand the problem

Reusing the KV-cache for stable prompt prefixes — the cheapest optimization most apps skip.

prompt-cachingcostlatencykv-cache
02

Attempt it yourself

Sketch your approach before reading the solution — that's what interviews test.

Nudge consolestandby

Stuck? Beam a request up — the console returns a conceptual nudge that guides your logic without spoiling the implementation.

03

Study the solution

Providers cache the internal computation (KV cache) for a prompt prefix they have seen recently, so repeated calls that share that prefix skip recomputing it — cached input tokens are billed at a large discount (up to ~90%) and time-to-first-token drops. To exploit it, structure prompts as stable-prefix-first (system p

Solution ready — 2 min read

Classified // press E to declassify

04

Read the code

Explicit cache breakpoints (Anthropic style)
await client.messages.create({
  model, max_tokens: 800,
  system: [
    { type: "text", text: BIG_STABLE_RUBRIC,          // 8k tokens, never changes
      cache_control: { type: "ephemeral" } },
  ],
  tools: TOOLS_SORTED_DETERMINISTICALLY,
  messages: [
    ...priorTurns,                       // growing cached prefix
    { role: "user", content: newQuestion },   // only this is fresh compute
  ],
});
05

Join the discussion

Discussion (0)

Sign in to join the discussion.

No responses yet. Be the first to share what you think.

Transmission complete // awaiting log

KEEP THE
STREAK ALIVE.

Dossier 37 of 80 decoded in the AI Engineering track. One more won't hurt.

Back to track