Data/ML/AIHigh signalEmergingAround since 2023

AI Evals

Also written as Evals, LLM Evaluation

Systematic tests that measure how well an AI system performs on a task, used to catch regressions and compare model or prompt changes objectively instead of eyeballing a few examples.

Think of it like

Like a standardized test given to every new hire doing the same job, so you can objectively compare performance instead of going on gut feeling from a few conversations.

Junior or senior?

Junior sounds like

Tinkers with prompts until one example looks good.

Senior sounds like

Can describe how they actually evaluated whether a prompt or model change improved things.

Ask them

“How did you evaluate whether a prompt or model change actually improved things?”