Why does adding more context make agents worse?
#What actually happens when you add a source
The intuition is that an agent with more information should give better answers. In practice, teams connect a fifth system and watch accuracy drop.
The mechanism is dilution. Similarity search works by finding the chunks closest to the question in vector space. When the corpus is small, the closest chunks are usually the right ones. As it grows, the space fills with text that is semantically near the question but contextually wrong: the same words, a different team, an older decision, a superseded policy.
This is measurable. Researchers at the University of Wyoming tracked a deployed system as its corpus scaled from 54 documents to 1,128. Retrieval accuracy fell from 75% to below 40%. Nothing about the model or the pipeline changed. Only the size of the haystack did. They called the failure mode vector search dilution.
#The three ways it gets worse at once
Precision falls. The top results are still close matches. They are just less likely to be the right ones, and nothing in the pipeline flags the difference.
Definitions start to conflict. At scale, the same term appears in multiple forms across the company. Finance and Growth define an active customer differently, and both are correct inside their own context. A retrieval system that ranks by similarity picks one at random and states it with confidence. This is the failure that gets caught latest, because the answer looks fine.
Cost and latency climb. An agent that cannot resolve a question keeps exploring: more retrievals, more pagination, more tokens. A well-scoped agent resolves a question in two or three tool calls. In our current benchmark, the same questions run through an agent without a reasoning layer underneath took 15 to 35 round trips to assemble an answer, and often gave up anyway. Every one of those costs tokens and time.
The same effect shows up in tool surface area, not just documents. Cloudflare found that exposing their full API as individual tools would consume roughly 1.17 million tokens of context before an agent did anything useful. Collapsing it to a compact interface cut that to about 1,000.
#Why a bigger context window doesn't fix it
Window size is the wrong lever. It governs how much text fits, not which text gets chosen. By the time the model sees the context, the selection has already happened, and the selection is what went wrong.
Pushing more candidates into a larger window makes it worse in a second way: the signal the model needs is now buried among near-misses that read as equally plausible.
#What actually fixes it
Reduce the search space before retrieval, not after.
If an engineer has worked in five repositories, the answer to their question is almost never in the other 4,995. The cheapest way to be precise is to never consider them. That means knowing who is asking, what they touch, and what the current question is adjacent to, then filtering before retrieval rather than ranking afterwards.
Three things have to be in place for that to work:
Entities resolved. The same service appears as billing-svc in Git, "Billing Service" in the wiki, BILL in the tracker and "the billing thing" in chat. Until those are one entity, narrowing by entity does nothing.
Terms defined. What your company means by a feature, a KPI, an active customer, encoded so a system can evaluate it rather than guess.
Intent understood. Who is asking and what their question sits next to, used to cut millions of candidates to the handful that could plausibly answer it.
Get this right and precision, latency and cost improve together. Skip it and they degrade together, which is why teams measuring only accuracy are confused by their token bill.
#How to tell if this is your problem
Measure retrieval precision separately from end-to-end answer quality. If you only track whether answers are good, you cannot tell whether the retrieval half or the generation half is broken.
Then track precision as the corpus grows. Dilution is gradual. It does not announce itself with an error, and it will not show up in an eval suite built when the index was a tenth of its current size.
#How does Uvi handle this?
Uvi narrows before it retrieves. The reasoning layer resolves entities across systems, holds the company's own definitions, and uses intent to cut the candidate set down before anything is fetched, so an agent gets the handful of things that bear on the question rather than everything that resembles it.
If you'd like to see how Uvi gives your AI agents accurate, current context, book a demo with us here.
Frequently asked questions
- Is this a model problem or a context problem?
- It is a context problem. The same model performs differently depending on what it is handed. If accuracy drops when you connect a new source, nothing about the model changed — only the candidate set did.
- Does a bigger context window fix it?
- No. A larger window changes how much text fits, not which text gets selected. The failure happens before the model sees anything: the wrong candidates were retrieved. A bigger window means you pay to send more of the wrong thing.
- Should we connect fewer systems, then?
- No. The answers really do live across systems, and disconnecting them trades one failure for another. The fix is narrowing the candidate set before retrieval runs, not shrinking the corpus.
- How do we tell whether this is happening to us?
- Measure retrieval precision separately from answer quality, and track it as the corpus grows. If answer quality falls while the index grows and the model stays the same, dilution is the most likely cause.