Data/ML/AI

LLM-as-a-Judge

Also written as LLM Judge, Model-Graded Evals, Auto-Graders

Using one language model to grade the output of another against a rubric, such as whether an answer is correct, grounded in the source documents, or on-brand. It lets teams evaluate thousands of responses without having people read every one.

Think of it like

Using a trained marker to grade exam scripts, with the head examiner spot-checking the marker's work.

Junior or senior?

The key question is whether they checked the judge against human ratings.

Senior sounds like

Measured how often the judge agreed with people and fixed the rubric when it didn't.

Ask them

“How did you know your LLM judge was grading correctly?”

Sounds like real experience

Describes comparing the judge's scores with human-labelled examples, and a bias or mistake they found and corrected.

Probe further if

Trusted the judge's scores as-is, with no comparison to human judgement.