Data/ML/AIBackend

Prompt Caching

Also written as Context Caching, KV Cache

Reusing the work a model has already done on the unchanging start of a prompt (long instructions, a document, a codebase) so later requests that share it are faster and much cheaper. Most major LLM APIs now offer it.

Think of it like

A teacher who has already read the textbook chapter doesn't re-read it before answering each student's question.

Junior or senior?

A practical cost and speed lever that separates people who have run LLM features at scale from people who've built demos.

Senior sounds like

Restructured prompts so the stable part came first, and can quote the cost or latency change.

Ask them

“Did you use prompt caching? How much did it change your costs or response time?”