A language model never sees the letters you type. Before your text reaches the model, a tokenizer cuts it into pieces called tokens and replaces each piece with a number. The model reads that list of numbers, predicts the next number, and the tokenizer turns the numbers back into text.
Once you understand tokens, several things that look strange about LLMs start to make sense: why a model can struggle to count the letters in a word, why a long document "does not fit", and why some languages cost more to process than others.
A tokenizer in one example
Take the sentence:
Packet sniffing is fun.A typical modern tokenizer might split it like this (each box is one token):
[Packet] [ sniff] [ing] [ is] [ fun] [.]Three things stand out:
- Spaces usually belong to the next word.
" is"with its leading space is a different token from"is"at the start of a line. - Common words are one token, rarer words are split. "is" and "fun" are common enough to have their own tokens. "sniffing" is split into a common stem and a common suffix.
- Punctuation is often its own token.
Each token is then mapped to an integer ID from the tokenizer's vocabulary. A vocabulary is a fixed list, often between 30,000 and 200,000 entries, learned before the model is trained.
You can try this yourself with the token counter, which runs a real tokenizer in your browser.
How tokenizers learn their vocabulary
Most current LLMs use a variant of byte-pair encoding (BPE). The idea is simple:
- Start with the smallest units: individual bytes.
- Count which pair of neighbouring units appears most often in a large training corpus.
- Merge that pair into a new unit and add it to the vocabulary.
- Repeat until the vocabulary reaches the target size.
Frequent sequences, like " the" or "ing", get merged early and become single tokens. Rare sequences stay split into several smaller tokens. Because the process starts from bytes, any text can be tokenized, even emoji or a language the tokenizer barely saw. It just costs more tokens.
Why token counts matter
Limits are measured in tokens
A model's context window is the maximum number of tokens it can consider at once: your instructions, the conversation so far, any retrieved documents, and the answer it writes. If you estimate in words, you will be wrong in a direction that is hard to predict.
Price and speed are measured in tokens
API providers bill per input token and per output token, and output tokens usually cost more. Generation speed is also reported in tokens per second. A prompt that is twice as many tokens costs roughly twice as much to send.
Some languages pay a "token tax"
Because vocabularies are learned mostly from English text, English words tend to be one token each. Text in Urdu, Arabic, Hindi and many other languages is often split into far more tokens for the same meaning. The same message can use several times the tokens, which means higher cost and less room in the context window. Paste the same sentence in English and Urdu into the token counter to see the difference on this site's tokenizer.
Tokens explain some odd model behaviour
- Counting letters is hard. If "strawberry" is two or three tokens, the model never directly sees the individual letters. It has to have learned the spelling as a fact.
- Arithmetic on long numbers is fragile. Numbers are often split into chunks of digits in inconsistent ways, which makes digit-by-digit reasoning harder.
- Tiny formatting changes change the input. An extra space or a different quote character produces different tokens, which can slightly change the output.
Rules of thumb
These are estimates for English text with modern tokenizers. Measure when it matters.
| Text | Rough token count |
|---|---|
| 1 short English word | 1 token |
| 1 English sentence (15 words) | about 20 tokens |
| 1 page of prose (500 words) | about 650 tokens |
| 100 lines of Python | often 800 to 1,500 tokens |
Counting tokens in Python
If you use OpenAI-family models, the tiktoken library exposes the same tokenizers they use:
import tiktoken
enc = tiktoken.get_encoding("o200k_base")
text = "Packet sniffing is fun."
ids = enc.encode(text)
print(len(ids)) # number of tokens
print([enc.decode([i]) for i in ids]) # each token as textOther model families publish their own tokenizers or offer a token-counting API endpoint. Use the one that matches your model.
Summary
- A token is a piece of text, usually a word or part of a word, mapped to an integer ID.
- Tokenizers are learned from data, so common English words are cheap and other scripts are often more expensive.
- Context limits, prices and speeds are all measured in tokens, so count them with the right tokenizer.
Next, see how tokens fill up a model's working memory in What is a context window?