Data/ML/AINew

AI Benchmarks

Also written as SWE-bench, MMLU, GPQA, HumanEval, LMArena, Chatbot Arena, SOTA

Standard tests used to compare AI models: SWE-bench for fixing real code issues, MMLU and GPQA for knowledge and expert reasoning, HumanEval for writing code, and LMArena for head-to-head human preference votes. 'SOTA' (state of the art) means best published score on one of them.

Think of it like

Standardised exams for models: useful for comparisons, but a high score doesn't prove someone can do your particular job.

Junior or senior?

Popular benchmarks get 'saturated' (every model scores near the top) or leak into training data, so serious teams build their own tests.

Senior sounds like

Has built a benchmark or evaluation for their own use case and can explain why public scores didn't predict real performance.

Ask them

“Did public benchmark scores match how the model performed on your own tasks? What did you measure instead?”