Skip to content

Learning path · Production RAG · 56

REFRAG

Retrieval compression pattern that keeps many candidate chunks as compact embeddings and expands only the ones the decoder needs back into tokens.

Why it matters

  • Cuts wasted tokens from stuffing oversized retrieval sets into the prompt.
  • Preserves broad recall while controlling latency and cost at decode time.
  • Fits production stacks already paying for large context windows.

Key ideas

  • Compressed retrieval
  • Selective expansion
  • Compute efficiency

Video

REFRAG-style pipelines keep most candidates as cheap embeddings and only expand a small set into readable tokens for the generator. That cuts the lost-in-the-middle tax and the bill for stuffing twenty near-duplicates into every call. Cap expanded chunks and tokens. Measure faithfulness before celebrating the cost win; compression that drops the supporting span is a silent regression.

Updated 2026-08-09 · Full learning path