Data/ML/AIAround since 2017

Transformer

The neural network architecture (introduced in 2017) behind nearly all modern LLMs, notable for processing entire sequences of text at once via a mechanism called 'attention.'

Think of it like

Like a reader who, instead of reading strictly one word at a time, glances at the whole sentence at once and weighs how much each word matters to every other word.

Junior or senior?

Junior sounds like

Has only called a transformer-based model through an API.

Senior sounds like

Can explain, at a high level, what the attention mechanism inside one is actually doing.

Ask them

“Can you explain, at a high level, what the 'attention' mechanism in a transformer is actually doing?”