Serving cached answers for similar (not identical) queries — big savings, real correctness risks.
01
01
Understand the problem
semantic-cachingcostcachingrisk
02
02
Attempt it yourself
Sketch your approach before reading the solution — that's what interviews test.
Stuck? AI Nudge Available
Get a conceptual hint to guide your logic without spoiling the final implementation.
03
03
Study the solution
The solution is waiting
Give it an honest attempt first — then compare your thinking with the full walkthrough.
04
04
Read the code
Guarded semantic cache lookup
async function cachedAnswer(q: string, ctx: Ctx) {
if (TRANSACTIONAL.test(q)) return llm(q, ctx); // never cache these
const key = { tenant: ctx.tenantId, promptV: PROMPT_VERSION, model: MODEL };
const hit = await semCache.nearest(await embed(q), { where: key, minSim: 0.95 });
if (hit && !expired(hit, ttlFor(hit.topic))) {
metrics.increment("semcache.hit");
return hit.answer;
}
const ans = await llm(q, ctx);
await semCache.put({ ...key, query: q, answer: ans });
return ans;
}05
05
Join the discussion
Discussion (0)
Sign in to join the discussion.
No responses yet. Be the first to share what you think.