Skip to content

Learning path · Embeddings & Representation · 35

Embeddings

Dense vector representations of text (or other modalities) where semantic similarity approximates geometric closeness.

Why it matters

  • Foundation of semantic search, RAG retrieval, and clustering.
  • Choice of embedding model affects recall on domain jargon.
  • Separate from generative LLM—you often use both.

Key ideas

  • Vector representations
  • Similarity search
  • Domain adaptation

Video

Embeddings turn a sentence into a vector so similar meanings sit nearby. That is how you find a passage without sharing keywords. Use an embedder trained on text like yours; legal and code often need their own. When you change the model, re-embed the corpus. Mixing two embedding versions in one index produces similarity scores that look precise and are nonsense.

Updated 2026-08-09 · Full learning path