LLM-as-a-Judge
Also written as LLM Judge, Model-Graded Evals, Auto-Graders
Using one language model to grade the output of another against a rubric, such as whether an answer is correct, grounded in the source documents, or on-brand. It lets teams evaluate thousands of responses without having people read every one.
Think of it like
Using a trained marker to grade exam scripts, with the head examiner spot-checking the marker's work.
Junior or senior?
The key question is whether they checked the judge against human ratings.
Senior sounds like
Measured how often the judge agreed with people and fixed the rubric when it didn't.
Ask them
“How did you know your LLM judge was grading correctly?”
Sounds like real experience
Describes comparing the judge's scores with human-labelled examples, and a bias or mistake they found and corrected.
Probe further if
Trusted the judge's scores as-is, with no comparison to human judgement.