hardAI Engineering

What is semantic caching and when is it safe to use?

954 views
01

Understand the problem

Serving cached answers for similar (not identical) queries — big savings, real correctness risks.

semantic-cachingcostcachingrisk
02

Attempt it yourself

Sketch your approach before reading the solution — that's what interviews test.

Stuck? AI Nudge Available

Get a conceptual hint to guide your logic without spoiling the final implementation.

03

Study the solution

The solution is waiting

Give it an honest attempt first — then compare your thinking with the full walkthrough.

04

Read the code

Guarded semantic cache lookup
async function cachedAnswer(q: string, ctx: Ctx) {
  if (TRANSACTIONAL.test(q)) return llm(q, ctx);          // never cache these

  const key = { tenant: ctx.tenantId, promptV: PROMPT_VERSION, model: MODEL };
  const hit = await semCache.nearest(await embed(q), { where: key, minSim: 0.95 });
  if (hit && !expired(hit, ttlFor(hit.topic))) {
    metrics.increment("semcache.hit");
    return hit.answer;
  }
  const ans = await llm(q, ctx);
  await semCache.put({ ...key, query: q, answer: ans });
  return ans;
}
05

Join the discussion

Discussion (0)

Sign in to join the discussion.

No responses yet. Be the first to share what you think.