Data/ML/AIBackendHigh signalEmergingAround since 2023

Time to First Token

Also written as TTFT, Streaming Response, Tokens Per Second

How quickly a model starts responding, as distinct from how long the full answer takes. Streaming the first words fast makes an application feel responsive even when total time is unchanged.

Think of it like

A waiter who acknowledges your order immediately versus one who disappears and returns twenty minutes later with everything at once.

Junior or senior?

Indicates someone who has shipped an AI product to real users rather than built a prototype.

Senior sounds like

Distinguishes perceived from total latency.

Ask them

“What did you do to make the AI feature feel fast, separately from making it actually faster?”