Learning path · Production RAG · 56
REFRAG
Retrieval compression pattern that keeps many candidate chunks as compact embeddings and expands only the ones the decoder needs back into tokens.
Why it matters
- Cuts wasted tokens from stuffing oversized retrieval sets into the prompt.
- Preserves broad recall while controlling latency and cost at decode time.
- Fits production stacks already paying for large context windows.
Key ideas
- Compressed retrieval
- Selective expansion
- Compute efficiency
Top resources
- 01DocsLlamaIndex
Understanding RAG
Why this resource. Treat REFRAG-class ideas as production RAG composition, not a magic model.
Covers in this concept
- compression
- retrieval loops
- 02PaperLewis et al.
Retrieval-Augmented Generation for Knowledge-Intensive NLP
Why this resource. Baseline you are departing from.
Covers in this concept
- naive RAG
REFRAG-style pipelines keep most candidates as cheap embeddings and only expand a small set into readable tokens for the generator. That cuts the lost-in-the-middle tax and the bill for stuffing twenty near-duplicates into every call. Cap expanded chunks and tokens. Measure faithfulness before celebrating the cost win; compression that drops the supporting span is a silent regression.
Updated 2026-08-09 · Full learning path