Learning path · Embeddings & Representation · 38
Chunking
Splitting documents into retrieval-sized pieces before embedding—balance context completeness against search precision.
Why it matters
- Bad chunking is the silent killer of RAG quality.
- Chunk size interacts with embedding model max length.
- Strategy depends on document structure and query type.
Key ideas
- Fixed-size splits
- Overlap windows
- Structure boundaries
Top resources
- 01DocsLlamaIndex
Node parsers / chunking
Why this resource. Node parsers: the practical chunking layer this concept describes.
Covers in this concept
- chunk size
- overlap
- nodes
- 02ArticleJina AI
Late chunking in long-context embedding models
Why this resource. When you should not chunk first—contrast for this concept.
Covers in this concept
- late chunking
Chunking decides what the retriever can return as atomic evidence. Fixed token windows are simple but split concepts mid-thought. Overlap reduces boundary cuts but increases storage. Compare semantic, late, and structure-aware strategies on your eval set—no universal best size. Always store metadata: source URL, heading path, page, and permissions for citation and filtering. Re-chunk when source formats change—PDF-to-Markdown migrations have invalidated more RAG systems than model upgrades ever did.
Updated 2026-08-09 · Full learning path