mediumAI Engineering

What is prompt caching and how does it reduce cost and latency?

599 views
01

Understand the problem

Reusing the KV-cache for stable prompt prefixes — the cheapest optimization most apps skip.

prompt-cachingcostlatencykv-cache
02

Attempt it yourself

Sketch your approach before reading the solution — that's what interviews test.

Stuck? AI Nudge Available

Get a conceptual hint to guide your logic without spoiling the final implementation.

03

Study the solution

The solution is waiting

Give it an honest attempt first — then compare your thinking with the full walkthrough.

04

Read the code

Explicit cache breakpoints (Anthropic style)
await client.messages.create({
  model, max_tokens: 800,
  system: [
    { type: "text", text: BIG_STABLE_RUBRIC,          // 8k tokens, never changes
      cache_control: { type: "ephemeral" } },
  ],
  tools: TOOLS_SORTED_DETERMINISTICALLY,
  messages: [
    ...priorTurns,                       // growing cached prefix
    { role: "user", content: newQuestion },   // only this is fresh compute
  ],
});
05

Join the discussion

Discussion (0)

Sign in to join the discussion.

No responses yet. Be the first to share what you think.