Data/ML/AIBackendHigh signalEmergingAround since 2023
Time to First Token
Also written as TTFT, Streaming Response, Tokens Per Second
How quickly a model starts responding, as distinct from how long the full answer takes. Streaming the first words fast makes an application feel responsive even when total time is unchanged.
Think of it like
A waiter who acknowledges your order immediately versus one who disappears and returns twenty minutes later with everything at once.
Junior or senior?
Indicates someone who has shipped an AI product to real users rather than built a prototype.
Senior sounds like
Distinguishes perceived from total latency.
Ask them
“What did you do to make the AI feature feel fast, separately from making it actually faster?”