Tokenization
Also written as Tokens
The process of breaking text into smaller pieces (tokens — often word fragments) that a language model actually processes and is priced/limited by, rather than reading raw characters or whole words.
Think of it like
Like a printing press setting text in individual pre-cast letter blocks rather than freehand writing — the model works with a fixed set of these blocks, not raw prose.
Junior or senior?
Junior sounds like
Has used an LLM API casually without hitting token limits.
Senior sounds like
Has had to optimize a prompt or pipeline specifically because of token limits or cost, and can describe the change.
Ask them
“Have you had to optimize a prompt or pipeline specifically because of token limits or token cost? What did you change?”