Packet To Sniff
LLM foundationsIntermediate

Why long context windows still miss things

A million-token context window does not mean a model uses every token well. Learn about the lost-in-the-middle effect, attention cost, and how to test it yourself.

By Packet To SniffPublished 3 min read

Helpful background: What is a context window in an LLM?.

On this page

Marketing pages quote context windows as a single number: 128K, 200K, 1M tokens. That number tells you how much text the model will accept. It does not tell you how well the model will use it. This article explains the gap and how to measure it for your own use case.

Accepting text is not the same as using it

Inside a transformer, every token can attend to every earlier token. The model learns during training how to spread that attention. Two practical problems follow:

  1. Compute grows quickly with length. In standard attention, the work for the attention step grows roughly with the square of the input length. Providers use many engineering tricks to make long inputs affordable, but longer requests are still slower and are billed per token.
  2. Training data is mostly short. Most documents a model learns from are far shorter than its maximum window. Models are extended to long contexts with extra training, but they have seen fewer examples of using information deep inside very long inputs.

The lost-in-the-middle finding

In 2023, Liu and colleagues published Lost in the Middle: How Language Models Use Long Contexts. They gave models a question plus a set of documents, where only one document contained the answer, and moved that document to different positions.

Accuracy formed a U shape: highest when the answer was near the beginning or the end of the input, and noticeably lower when it was in the middle. For some settings, putting the answer in the middle did worse than giving the model no documents at all.

Output
accuracy
  high |  *                         *
       |    *                     *
       |       *               *
   low |           *   *   *
       +------------------------------
         start      middle        end
             position of the answer

The shape of the result, simplified. Exact numbers depend on the model and task.

Newer long-context models do much better on simple tests, such as finding a single planted sentence (often called a "needle in a haystack" test). But finding one sentence is the easy case. Tasks that need the model to connect several facts spread across a long input remain harder, and results differ a lot between models.

What this means when you build

  • Put the most important material near the start or end. Many teams place retrieved sources near the end, right before the question, and keep instructions at the start.
  • Retrieve, then rank. Sending the five best passages usually beats sending fifty passable ones. That is the core argument for RAG even when a model's window is huge.
  • Keep each passage self-contained. Good chunking means a passage still makes sense when it lands in the middle.
  • Budget for cost. A request that sends 500,000 tokens pays for 500,000 tokens every time it runs.

Test it yourself

Do not trust a benchmark chart for your use case. A small experiment tells you more:

  1. Write 20 questions whose answers each appear in one known passage of a long document.
  2. For each question, build prompts where that passage sits at 10%, 50% and 90% of the input.
  3. Ask the model and score each answer as correct or not.
  4. Compare accuracy by position.
Python
def build_prompt(filler_passages, answer_passage, position, question):
    """Insert the answer passage at a relative position (0.0 to 1.0)."""
    idx = int(position * len(filler_passages))
    passages = filler_passages[:idx] + [answer_passage] + filler_passages[idx:]
    context = "\n\n".join(passages)
    return f"Context:\n{context}\n\nQuestion: {question}\nAnswer briefly."
 
positions = [0.1, 0.5, 0.9]
# For each question and position: call your model, score the answer,
# then compute accuracy per position.

Long context vs retrieval

Long context Retrieval (RAG)
Setup effort Low: paste everything Higher: chunk, index, retrieve
Cost per question High for large inputs Low: only relevant chunks
Works when the answer needs the whole document Better Worse
Shows which source was used Hard Natural: citations per chunk
Knowledge updates Re-send new text Re-index changed documents

They are not rivals. Many good systems retrieve the right sections, then use a comfortably large window to include them with their surrounding context.

Summary

  • The advertised window is how much text a model accepts, not how well it uses it.
  • Models have been shown to recall information in the middle of long inputs less reliably.
  • Put key material near the ends, send less but better text, and measure position effects on your own task.

Frequently asked questions

What is the lost in the middle effect?

It is the finding, reported by Liu and colleagues in 2023, that language models answered questions more accurately when the relevant passage was near the beginning or end of a long input, and less accurately when it sat in the middle.

Have newer models fixed this?

Newer long-context models perform much better on simple retrieval tests, but results still vary by model, task and input length. Harder tasks, such as combining several facts spread across a long document, remain more difficult. Test the model you actually use.

Continue learning

Tags

  • What is a token in an LLM?

    A token is the unit of text a language model reads and writes. Learn how tokenizers split text, why token counts differ by language, and why it matters for cost.

    LLM foundationsBeginner4 min
  • What are embeddings? Meaning as numbers, explained

    Embeddings turn text into lists of numbers so that similar meanings land close together. Learn how they work, how similarity is measured, and where they fail.

    LLM foundationsBeginner3 min
  • Cross-validation explained: k-fold, stratified and grouped

    Cross-validation gives a more reliable performance estimate than one split by rotating the validation fold. Learn k-fold, stratified, group and time-series variants.

    Machine learning methodsIntermediate3 min