Skip to content

Learning path · Evaluation & Quality · 73

LLM as Judge

Using a strong model to score another model's outputs against rubrics—relevance, safety, coherence.

Why it matters

  • Scales eval beyond manual review for fast iteration.
  • Biases exist—judges favor verbose or self-similar styles.
  • Calibrate against human labels regularly.

Key ideas

  • Rubric prompts
  • Pairwise comparison
  • Judge bias

Video

LLM judges apply a rubric: score 1 to 5 whether each claim is supported by the passage. That lets you sweep prompt variants overnight. Rotate judge models, keep a human-labeled gold set, and do not promote on judge scores alone. Judge drift has quietly invalidated promotion decisions.

Updated 2026-08-09 · Full learning path