Learning path · Evaluation & Quality · 73
LLM as Judge
Using a strong model to score another model's outputs against rubrics—relevance, safety, coherence.
Why it matters
- Scales eval beyond manual review for fast iteration.
- Biases exist—judges favor verbose or self-similar styles.
- Calibrate against human labels regularly.
Key ideas
- Rubric prompts
- Pairwise comparison
- Judge bias
Top resources
- 01PaperLiu et al.
G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
Why this resource. The G-Eval protocol for rubric-based LLM judges.
Covers in this concept
- form-filling
- chain-of-thought judge
- 02DocsExploding Gradients
RAGAS
Why this resource. Where judge metrics show up in a RAG eval harness.
Covers in this concept
- answer quality
LLM judges apply a rubric: score 1 to 5 whether each claim is supported by the passage. That lets you sweep prompt variants overnight. Rotate judge models, keep a human-labeled gold set, and do not promote on judge scores alone. Judge drift has quietly invalidated promotion decisions.
Updated 2026-08-09 · Full learning path