Learning path · Evaluation & Quality · 72
Hallucination
Model outputs that sound plausible but are factually unsupported or contradict provided evidence.
Why it matters
- Top risk in customer-facing and compliance workflows.
- RAG without faithfulness checks can increase confident errors.
- Detection blends automated metrics and human audit.
Key ideas
- Unsupported claims
- Confident tone
- Faithfulness testing
Hallucinations thrive when questions exceed context, retrieval misses, or prompts forbid "I don't know." Mitigate with grounding requirements, retrieval confidence thresholds, and faithfulness evals. Monitor citation click-through and support escalations as lagging indicators. Train support staff that fluent ≠ verified. Track hallucination rate alongside business metrics—support deflection means nothing if escalations spike due to wrong policies. Treat hallucination monitoring as a production checklist item, not a research curiosity, before you scale traffic or spend. Ship only after eval gates pass on representative production failures.
Updated 2026-08-09 · Full learning path