A million tokens can still lose the right document

Suppose you ask, "Which report explains this failure?" and the answer is somewhere in ten thousand files. The usual approach is to search first. Retrieval-augmented generation, or RAG, finds likely files and gives just those to a language model to read. The search might match keywords or compare numerical representations of the question and documents. If its first picks are noisy, a reranker can read that smaller shortlist and reorder it. An agent can also search again, try a different query, and follow a promising lead. None of these methods is magic. A missed document cannot help the final answer, and a retrieved one can still be misunderstood.
That search step is useful, but it can lose information. A chunk is a small piece cut from a longer document. If a rule lands in one chunk and its exception in another, returning only the rule changes what the model sees. A question might also need a fact from one report and a correction from another. Unless both are retrieved, the model cannot reason from both. Rerankers improve the order of the candidates they receive, but they cannot rank a document that never reached their shortlist. An agent that searches again can recover some misses, although its next query still has to find the missing evidence. For an answer grounded in those files, what retrieval brings back limits what the model can support.
A million-token context window suggests another route: skip the search system and put the whole collection in front of the model. A token is a small unit of text; the context window is how much text the model can take in at once. Why make a shortlist if the model can see every file? It removes the need to trust a shortlist to preserve every useful connection. That sounds like it could replace the retrieval step in RAG. But "in the input" is not the same as "found." The model still has to choose the right document from thousands of distractions.
Siddharth Gollapudi, Nilesh Gupta, Prasann Singhal, and Sewon Min test that specific job in Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale. They are not testing whether an end-to-end RAG assistant gives a good answer, or whether an agent with repeated searches can recover a miss. Their task is narrower: given a question and a collection already in context, can a model name the relevant document, even as the input approaches a million tokens? To test it, they build BlockSearch, a language model trained to find a document in the collection it has been given. Following how it makes that choice helps explain why seeing every file is not enough.
A needle that is visible but loses its weight
BlockSearch turns that broad question into a simple job: point to the right file. The authors give each document a short identifying code, so the model can return that code instead of copying the document or writing a full answer. If the failure is explained in Report A, its job is to choose Report A. Unlike the search-first approach, it receives the question and the full collection together, so its choice has to come from a much longer input.
At smaller collection sizes, the model can often select the right document. As the number of documents grows, its answer gets less reliable. The paper traces one reason through attention, the mechanism that lets a model give different parts of its input different weight. Imagine the correct document retains the highest individual relevance score. Every distracting document still contributes a little to the total used to turn those scores into weights. With enough distractors, the correct document's share can become too small for the model's final answer to reflect it. The authors call that attention dilution. It is not simply a case of the document being outside the input.

What the experiment actually asks
To see whether this loss of influence also changes which file gets picked, the authors test BlockSearch on collections with known correct choices. They use three document-search test sets: Natural Questions, MS MARCO, and HotpotQA. Each supplies questions and documents marked as relevant. They put those documents into a collection alongside competing documents, including texts that a search system finds plausible but that are not the marked answer. Then they grow the collection while keeping the relevant evidence available. BlockSearch must return a document identifier, not write an answer to the question.
Because BlockSearch is choosing a file, the measure is whether its first pick is a marked relevant document. The paper calls this recall at rank one. In tasks with several acceptable documents, finding any one can count as success. That is a retrieval test, not a test of whether an assistant can combine all necessary evidence and answer correctly. The unmodified model's first-pick success rate falls as the collection grows. The next experiments ask whether changing its attention can recover that lost performance.
Two repairs, and why one looks like RAG
The authors try two repairs to that attention problem, each aimed at stopping useful evidence from being overwhelmed by the rest of the collection. The weighting repair makes the gap between high and low token scores sharper as the collection grows. For our failure-report question, the aim is to let the strong signal from the relevant passage stand out more against weaker signals elsewhere. It does not pick a document shortlist; it changes how much influence the tokens receive. The routing repair does pick a shortlist: earlier model stages score the documents, and later dense attention works only on the highest-ranked ones. A layer is one stage of the model's computation. The selection happens within that computation, rather than through a separate search index.

That shortlist brings us back to the search-first approach we started with. Is routing basically RAG again? Yes: structurally, it is retrieve-then-read. The authors explicitly say it brings a "retrieve-then-read" structure back inside the model. A usual RAG pipeline selects documents through a separate search step before an answering model reads them. Here, the model's own earlier layers supply the selection scores. The similarity is real: both trust a shortlist to preserve the useful evidence. The difference is where the shortlist comes from, not that this method has somehow stopped doing retrieval.
Moving the shortlist inside the model changes who makes the selection, not whether a selection is made. Routing avoids handing selection to a separate retriever in this experiment, but it does not remove selection errors. Later dense attention cannot return to the tokens of a document excluded from its shortlist. The authors can inspect the internal scores and test the shortlist size, so this is not an entirely invisible decision either. Their results do not settle whether it is better for a production assistant than a search pipeline whose selection and reading steps can be changed separately.
What improved, in which setting?
The comparison below asks whether those two changes help BlockSearch choose the right file. It is one specific MS MARCO experiment, not a result for all long-context tasks. Its collection contains 10,000 documents, taking about 900,000 input tokens. The test uses 400 questions. For each question, success means the model's first document choice is marked relevant. On that same test, the unmodified BlockSearch gets 0.2% first-pick recall; with the sharper weighting and routing together, it gets 20.5%. The percentages are the reported retrieval rates, not answer accuracy or the share of a document the model understood.

That is a large recovery from almost no successful first picks, but roughly four in five questions still miss in this setting. The broader tests show gains on other collection sizes and tasks too. A dense retriever, which searches through numerical representations of documents, remains a useful comparison: on this MS MARCO test it gets 20.2%, close to the repaired model. On LIMIT, a benchmark built so word-matching questions trip up semantic search, BlockSearch does better than that baseline. None of this is a comparison with a complete RAG assistant using rerankers or repeated agent searches. It also does not show that a whole archive in context is cheaper, safer, or better maintained than an index.
What if most of the documents matter?
Suppose an answer needs facts from 1,000 reports, rather than one report hidden among distractions. You might be comparing failures across all of them, so finding just one good report would not be enough. But the routing setup tested in this paper keeps only 256 documents. If all 1,000 are needed, 744 are left out of the later reading step. The authors chose that limit; the model does not increase it when the question needs more evidence.

Which reports stay? Different parts of the model, called attention heads, score how well each report matches the question, those scores are added together, and the top 256 stay. That is a ranking of the model's best guesses, not a count of how many reports the answer actually needs. Among the reports it keeps, some passages still get more attention than others. Two equally useful reports need not get equal attention, because the model's scores can differ from their real usefulness.
The earlier stages have seen the full collection, but the later reading step cannot go back to the excluded reports' text. The weighting change alone keeps all the documents, yet that does not prove it can combine facts from all of them correctly. The paper tests finding a relevant document, not writing an answer that needs most or all of the collection. So these results do not tell us whether either repair can handle that larger job.
An idea for chat history
Reading this made me think about chat history. We can run into the same problem there: an earlier message might be left out when we retrieve a few turns, or it might still be in the context and get overlooked. I had an idea for a chat assistant: keep the conversation together, then search the larger library for whatever else the answer needs. The paper does not test this setup, but its warning about losing useful evidence seems worth checking here too.
Imagine a support chat where I say, "I've already tried resetting it," and later ask, "What should I try next?" If the assistant retrieves only that last question, it could suggest the reset again. I would keep the conversation in context while it fits, so the earlier attempt stays alongside the new question. Then I would retrieve the relevant instructions from the support library. The conversation tells it what has happened; the library supplies the instructions. There is no need to put the whole library in every request.

To avoid processing the same opening text from scratch each time, I would also use context caching. The unchanged conversation prefix and the assistant's instructions can be reused, with new turns and retrieved evidence added after them. A prefix just means the shared text at the start of each request. Caching can lower the cost of processing that reused text and sometimes reduce delay, although explicit caches can have storage charges. It does not stop the conversation from growing.
Keeping the turns together would avoid dropping one just because it missed a retrieval shortlist. It would not guarantee that the assistant pays attention to it. That is the part I would want to test: does it remember the failed reset and suggest something else, even after more messages have been added? Long chats may still need summaries or retrieval, and caching does not fix attention dilution. This is my proposed use of the idea, not a result from the paper.
What I would test next
That leaves a practical question: how would I check whether extra context helps or hides the evidence in my own system? I would keep the question and relevant document fixed, then add unrelated text, documents on the same topic, and near duplicates differing in one important fact. Does the right document survive the routing shortlist, and does the final answer pick it? Then change the supporting passage and check whether the answer updates and cites the new evidence. These are proposed tests, not results reported in this paper.
The larger input is only an opportunity to find the right evidence. This paper shows why the model can have that evidence in front of it and still miss it.
Paper: Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale
Sources for the practical examples: Context caching overview; Mitigating Lost-in-Retrieval Problems in Retrieval Augmented Multi-Hop Question Answering.