Learning path · Retrieval & Ranking · 43
Semantic Search
Finding documents by meaning similarity between query and corpus embeddings rather than exact keyword match.
Why it matters
- Captures paraphrases and conceptual questions keywords miss.
- Core retrieval stage before RAG generation.
- Fails on rare proper nouns without hybrid lexical backup.
Key ideas
- Query embedding
- Top-K retrieval
- Similarity thresholds
Top resources
Semantic search embeds the question, pulls nearest neighbours, and hands them to a reranker or the LLM. Too few hits miss evidence; too many waste tokens. Log zero-hit queries: they usually mean a stale index or a missing glossary term. Filter by ACL on every query so embeddings never leak another tenant's docs.
Updated 2026-08-09 · Full learning path