Packet To Sniff

What are embeddings? Meaning as numbers, explained

Embeddings turn text into lists of numbers so that similar meanings land close together. Learn how they work, how similarity is measured, and where they fail.

By Packet To SniffPublished 3 min read

Helpful background: What is a token in an LLM?.

On this page

Search engines used to work mostly by matching words. If you searched for "car" you would miss a page that only said "automobile". Embeddings fix this by representing meaning as numbers, so that "car" and "automobile" end up close together even though they share no letters.

Embeddings are the engine under semantic search, recommendation systems, clustering, and the retrieval step of RAG.

From text to a vector

An embedding model takes text and returns a fixed-length list of numbers:

Output
"How do I reset my router?"
  -> [0.021, -0.113, 0.087, ..., 0.004]   (for example, 384 numbers)

Every input, whether a single word or a full paragraph, produces a vector of the same length. The individual numbers are not meaningful on their own. What matters is how vectors relate to each other.

The model is trained on huge numbers of pairs of texts that should be close (a question and its answer, a title and its article) and pairs that should be far apart. Over training, it learns to place related meanings near each other.

Similar meaning, similar direction

Picture each vector as an arrow from the origin. Texts about networking point roughly one way; texts about cooking point another. We usually compare two arrows with cosine similarity, which measures the angle between them:

  • 1.0: same direction, very similar meaning
  • around 0: unrelated
  • negative: pointing away from each other (less common in practice with modern text embeddings)

The formula divides the dot product by the lengths of the two vectors:

Output
cosine(a, b) = (a · b) / (|a| × |b|)

In code:

Python
import numpy as np
 
def cosine(a, b):
    a, b = np.asarray(a), np.asarray(b)
    return float(a @ b / (np.linalg.norm(a) * np.linalg.norm(b)))
 
print(cosine([1, 0], [1, 0]))    # 1.0  same direction
print(cosine([1, 0], [0, 1]))    # 0.0  unrelated
print(cosine([1, 1], [2, 2]))    # 1.0  length does not matter, direction does

A small, real example

Using the open-source sentence-transformers library and the small all-MiniLM-L6-v2 model, which produces 384-dimensional vectors:

Python
from sentence_transformers import SentenceTransformer, util
 
model = SentenceTransformer("all-MiniLM-L6-v2")
texts = [
    "How do I reset my Wi-Fi router?",
    "Steps to restart a wireless access point",
    "Best recipe for chicken karahi",
]
vectors = model.encode(texts, normalize_embeddings=True)
print(util.cos_sim(vectors[0], vectors[1]))  # high: same topic, different words
print(util.cos_sim(vectors[0], vectors[2]))  # low: unrelated

The first two sentences share almost no words, yet their similarity is high. That is the whole point.

Where embeddings are used

Use How embeddings help
Semantic search Find passages whose meaning matches a query
RAG retrieval Pick the chunks to put into the context window
Deduplication Spot near-duplicate documents or tickets
Clustering Group support tickets or emails by topic
Classification Use vectors as features for a simple classifier

Where embeddings fail

Embeddings are powerful but not magic. Know their weak spots:

  1. Exact identifiers. Error codes, CVE numbers, IP addresses and function names are often better matched by keyword search. This is why many systems use hybrid search.
  2. Negation and small words. "Allowed" and "not allowed" can land surprisingly close together.
  3. Domain vocabulary. A general model may not separate specialised terms well. Test with your own documents.
  4. Length mismatch. A short question and a long passage are different kinds of text. Some models are trained for this "asymmetric" case and expect a prefix such as query: or passage:. Read the model card.
  5. Model changes. Vectors from different models are incompatible. Changing models means re-embedding everything.

Summary

  • An embedding is a vector that represents meaning; similar texts get vectors pointing in similar directions.
  • Cosine similarity compares direction and is the usual way to score closeness.
  • Embeddings find paraphrases well but can miss exact identifiers and negation, so test on your own data and consider hybrid search.

Frequently asked questions

What is the difference between a token and an embedding?

A token is a piece of text mapped to an integer ID. An embedding is a vector of many decimal numbers that represents meaning. Inside a model, each token ID is looked up as a vector, and embedding models combine those into one vector for a whole sentence or passage.

Can I compare embeddings from two different models?

No. Each embedding model defines its own space. Vectors from different models, or even different versions of the same model, are not comparable. If you change the embedding model, re-embed all your documents.

How many dimensions should an embedding have?

It depends on the model. Common models produce vectors with a few hundred to a few thousand dimensions. More dimensions can capture more detail but cost more storage and compute. Pick a model by testing retrieval quality on your own data.

Continue learning

Tags

  • How to chunk documents for RAG

    Chunking decides what your RAG system can retrieve. Compare fixed-size, recursive, heading-aware and semantic chunking, and learn how to pick size and overlap.

    Retrieval-augmented generationIntermediate3 min
  • What is a context window in an LLM?

    The context window is the maximum number of tokens a language model can consider at once. Learn what fills it, what happens when it overflows, and how to budget it.

    LLM foundationsBeginner3 min
  • Why long context windows still miss things

    A million-token context window does not mean a model uses every token well. Learn about the lost-in-the-middle effect, attention cost, and how to test it yourself.

    LLM foundationsIntermediate3 min