Data/ML/AIHigh signalEmergingAround since 2023
AI Evals
Also written as Evals, LLM Evaluation
Systematic tests that measure how well an AI system performs on a task, used to catch regressions and compare model or prompt changes objectively instead of eyeballing a few examples.
Think of it like
Like a standardized test given to every new hire doing the same job, so you can objectively compare performance instead of going on gut feeling from a few conversations.
Junior or senior?
Junior sounds like
Tinkers with prompts until one example looks good.
Senior sounds like
Can describe how they actually evaluated whether a prompt or model change improved things.
Ask them
“How did you evaluate whether a prompt or model change actually improved things?”