LLM Inference Optimization
Also written as Inference Optimization, Speculative Decoding, Continuous Batching, PagedAttention, FlashAttention, TensorRT-LLM, SGLang, llama.cpp
Making AI models answer faster and more cheaply on the same hardware. Techniques include serving many requests together (continuous batching), using a small model to draft words a big model then checks (speculative decoding), managing memory more efficiently, and compressing the model. vLLM, SGLang, TensorRT-LLM and llama.cpp are common tools.
Think of it like
Running a restaurant kitchen so the same chefs serve twice as many diners without the food getting worse.
Junior or senior?
Scarce and valuable wherever a company runs its own models at scale.
Senior sounds like
Quotes before-and-after numbers — speed, throughput, cost per request — and the quality checks they ran.
Ask them
“What was the biggest speed or cost improvement you made, and how did you check answer quality didn't drop?”