Marketing pages quote context windows as a single number: 128K, 200K, 1M tokens. That number tells you how much text the model will accept. It does not tell you how well the model will use it. This article explains the gap and how to measure it for your own use case.
Accepting text is not the same as using it
Inside a transformer, every token can attend to every earlier token. The model learns during training how to spread that attention. Two practical problems follow:
- Compute grows quickly with length. In standard attention, the work for the attention step grows roughly with the square of the input length. Providers use many engineering tricks to make long inputs affordable, but longer requests are still slower and are billed per token.
- Training data is mostly short. Most documents a model learns from are far shorter than its maximum window. Models are extended to long contexts with extra training, but they have seen fewer examples of using information deep inside very long inputs.
The lost-in-the-middle finding
In 2023, Liu and colleagues published Lost in the Middle: How Language Models Use Long Contexts. They gave models a question plus a set of documents, where only one document contained the answer, and moved that document to different positions.
Accuracy formed a U shape: highest when the answer was near the beginning or the end of the input, and noticeably lower when it was in the middle. For some settings, putting the answer in the middle did worse than giving the model no documents at all.
accuracy
high | * *
| * *
| * *
low | * * *
+------------------------------
start middle end
position of the answerThe shape of the result, simplified. Exact numbers depend on the model and task.
Newer long-context models do much better on simple tests, such as finding a single planted sentence (often called a "needle in a haystack" test). But finding one sentence is the easy case. Tasks that need the model to connect several facts spread across a long input remain harder, and results differ a lot between models.
What this means when you build
- Put the most important material near the start or end. Many teams place retrieved sources near the end, right before the question, and keep instructions at the start.
- Retrieve, then rank. Sending the five best passages usually beats sending fifty passable ones. That is the core argument for RAG even when a model's window is huge.
- Keep each passage self-contained. Good chunking means a passage still makes sense when it lands in the middle.
- Budget for cost. A request that sends 500,000 tokens pays for 500,000 tokens every time it runs.
Test it yourself
Do not trust a benchmark chart for your use case. A small experiment tells you more:
- Write 20 questions whose answers each appear in one known passage of a long document.
- For each question, build prompts where that passage sits at 10%, 50% and 90% of the input.
- Ask the model and score each answer as correct or not.
- Compare accuracy by position.
def build_prompt(filler_passages, answer_passage, position, question):
"""Insert the answer passage at a relative position (0.0 to 1.0)."""
idx = int(position * len(filler_passages))
passages = filler_passages[:idx] + [answer_passage] + filler_passages[idx:]
context = "\n\n".join(passages)
return f"Context:\n{context}\n\nQuestion: {question}\nAnswer briefly."
positions = [0.1, 0.5, 0.9]
# For each question and position: call your model, score the answer,
# then compute accuracy per position.Long context vs retrieval
| Long context | Retrieval (RAG) | |
|---|---|---|
| Setup effort | Low: paste everything | Higher: chunk, index, retrieve |
| Cost per question | High for large inputs | Low: only relevant chunks |
| Works when the answer needs the whole document | Better | Worse |
| Shows which source was used | Hard | Natural: citations per chunk |
| Knowledge updates | Re-send new text | Re-index changed documents |
They are not rivals. Many good systems retrieve the right sections, then use a comfortably large window to include them with their surrounding context.
Summary
- The advertised window is how much text a model accepts, not how well it uses it.
- Models have been shown to recall information in the middle of long inputs less reliably.
- Put key material near the ends, send less but better text, and measure position effects on your own task.