Search engines used to work mostly by matching words. If you searched for "car" you would miss a page that only said "automobile". Embeddings fix this by representing meaning as numbers, so that "car" and "automobile" end up close together even though they share no letters.
Embeddings are the engine under semantic search, recommendation systems, clustering, and the retrieval step of RAG.
From text to a vector
An embedding model takes text and returns a fixed-length list of numbers:
"How do I reset my router?"
-> [0.021, -0.113, 0.087, ..., 0.004] (for example, 384 numbers)Every input, whether a single word or a full paragraph, produces a vector of the same length. The individual numbers are not meaningful on their own. What matters is how vectors relate to each other.
The model is trained on huge numbers of pairs of texts that should be close (a question and its answer, a title and its article) and pairs that should be far apart. Over training, it learns to place related meanings near each other.
Similar meaning, similar direction
Picture each vector as an arrow from the origin. Texts about networking point roughly one way; texts about cooking point another. We usually compare two arrows with cosine similarity, which measures the angle between them:
- 1.0: same direction, very similar meaning
- around 0: unrelated
- negative: pointing away from each other (less common in practice with modern text embeddings)
The formula divides the dot product by the lengths of the two vectors:
cosine(a, b) = (a · b) / (|a| × |b|)In code:
import numpy as np
def cosine(a, b):
a, b = np.asarray(a), np.asarray(b)
return float(a @ b / (np.linalg.norm(a) * np.linalg.norm(b)))
print(cosine([1, 0], [1, 0])) # 1.0 same direction
print(cosine([1, 0], [0, 1])) # 0.0 unrelated
print(cosine([1, 1], [2, 2])) # 1.0 length does not matter, direction doesA small, real example
Using the open-source sentence-transformers library and the small all-MiniLM-L6-v2 model, which produces 384-dimensional vectors:
from sentence_transformers import SentenceTransformer, util
model = SentenceTransformer("all-MiniLM-L6-v2")
texts = [
"How do I reset my Wi-Fi router?",
"Steps to restart a wireless access point",
"Best recipe for chicken karahi",
]
vectors = model.encode(texts, normalize_embeddings=True)
print(util.cos_sim(vectors[0], vectors[1])) # high: same topic, different words
print(util.cos_sim(vectors[0], vectors[2])) # low: unrelatedThe first two sentences share almost no words, yet their similarity is high. That is the whole point.
Where embeddings are used
| Use | How embeddings help |
|---|---|
| Semantic search | Find passages whose meaning matches a query |
| RAG retrieval | Pick the chunks to put into the context window |
| Deduplication | Spot near-duplicate documents or tickets |
| Clustering | Group support tickets or emails by topic |
| Classification | Use vectors as features for a simple classifier |
Where embeddings fail
Embeddings are powerful but not magic. Know their weak spots:
- Exact identifiers. Error codes, CVE numbers, IP addresses and function names are often better matched by keyword search. This is why many systems use hybrid search.
- Negation and small words. "Allowed" and "not allowed" can land surprisingly close together.
- Domain vocabulary. A general model may not separate specialised terms well. Test with your own documents.
- Length mismatch. A short question and a long passage are different kinds of text. Some models are trained for this "asymmetric" case and expect a prefix such as
query:orpassage:. Read the model card. - Model changes. Vectors from different models are incompatible. Changing models means re-embedding everything.
Summary
- An embedding is a vector that represents meaning; similar texts get vectors pointing in similar directions.
- Cosine similarity compares direction and is the usual way to score closeness.
- Embeddings find paraphrases well but can miss exact identifiers and negation, so test on your own data and consider hybrid search.