Embeddings and Semantic Search: How RAG Finds Meaning
An embedding is a vector — a list of numbers — that represents a text's meaning, such that texts with similar meaning get similar vectors, sitting close together in a high-dimensional space.
This is what lets RAG do semantic search: finding passages by meaning rather than exact keywords, so a question and the answer that means the same thing match even when they share no words.
From meaning to numbers
An embedding model takes a piece of text and returns a vector — often several hundred or a few thousand numbers. The crucial property is that texts with similar meaning get similar vectors, close together in the space, while unrelated texts land far apart. "How do I reset my password?" and "I forgot my login credentials" share almost no words but mean nearly the same thing, and a good embedding places them near each other.
Why it beats keyword search
Traditional search matches words. Ask about "cars" and if your document says "automobiles," keyword search misses it. Embeddings match meaning, so the synonym is found. This semantic search is the capability that makes RAG work on real questions, which rarely use the exact words of the source.
It's not magic, though. Embeddings capture topical similarity well and precise logical distinctions poorly, which is why strong systems add keyword search back alongside them — a combination called hybrid search.
Keyword search matches words.
Embeddings match meaning.
Choosing an embedding model
- Quality — how well it separates relevant from irrelevant on text like yours. Leaderboards guide you; your own data decides.
- Dimension — bigger vectors capture more but cost more to store and search. Often a mid-size model is the sweet spot.
- Domain fit — a model trained on general web text may struggle with legal, medical, or code. Match the model to your material.
- Consistency — you must embed chunks and queries with the same model. Change it and you must re-embed everything.
The query and the document must match
One rule causes more silent failures than any other: the question and the chunks must be embedded by the same model, into the same space, or the distances are meaningless. If you upgrade your embedding model, you can't just embed new queries — you must re-index the entire corpus. Treat the embedding model as a foundation you change deliberately, not casually.
Embeddings are not the model
A common confusion: the embedding model that powers retrieval is separate from the language model that writes the answer. They're different models doing different jobs — one finds text, the other reads it. Choosing them is two independent decisions, and the embedding model is the one that quietly determines whether retrieval works at all.