Data/ML/AIDevOps/CloudHigh signalAround since 2020
Model Serving
Also written as Inference Server, vLLM, Inference Endpoint
The infrastructure that runs a model and answers requests — handling batching, GPU memory and concurrency so that responses stay fast under load.
Think of it like
The difference between owning a professional oven and running a restaurant service with it.
Junior or senior?
Strong signal for anyone self-hosting models rather than calling an API.
Senior sounds like
Talks about batching, GPU utilisation and cost per request.
Ask them
“How did you keep latency acceptable as concurrent requests went up?”