A million tokens of input is a frontier feature now: Claude Opus 5 and Sonnet 5, GPT-6-astra, Gemini. Not the cheap tiers, and not Claude Haiku 4.5, which stops at 200K.
Anthropic has also removed the reason to be sparing with it. Its pricing docs put that in a parenthesis: "A 900k-token request is billed at the same per-token rate as a 9k-token request." (opens in a new tab) Google puts a million tokens at roughly eight average-length English novels or 50,000 lines of code (opens in a new tab).
So the obvious move is to stop curating. No chunking, no retrieval, no knowledge base to maintain. Paste the company in and ask the question.
Nobody who ships agents actually does that, the vendors included. The reason turned out to be narrower than "long context is unreliable," and working it out moved me somewhere I wasn't expecting.
The people selling the window pay not to fill it
Anthropic's own docs tell you to reach for a subagent (opens in a new tab) when "a side task would flood your main conversation... the subagent does that work in its own context and returns only the summary." Its cost docs (opens in a new tab) price that habit: agent teams "use approximately 7x more tokens than standard sessions... because each teammate maintains its own context window." Its multi-agent research post (opens in a new tab) puts multi-agent systems at roughly 15x the tokens of a chat, against about 4x for a single agent.
Read those together. People are paying seven to fifteen times the tokens to avoid putting everything in one window that carries no premium for being full.
The bill buys something, to be fair: Anthropic reports that multi-agent setup beating single-agent Opus 4 by 90.2%, and some of that will be parallel search rather than context hygiene. But the reason its own docs give for reaching for a subagent is the flood.
That's a revealed preference, and it needed no benchmark to establish. Sourcegraph (opens in a new tab) measured the same thing from the other end. On identical tasks, its agents did worse given a 100K-token codebase summary than given 5K of targeted retrieval. More relevant material, worse outcome.
The vendors also say it in plain text. Anthropic's context windows page (opens in a new tab), on the same site advertising a million tokens: "more context isn't automatically better. As token count grows, accuracy and recall degrade, a phenomenon known as context rot. This makes curating what's in context just as important as how much space is available." Google concedes the same (opens in a new tab) for multi-needle work: "the model does not perform with the same accuracy."
Then a number that does more work than any of the prose. In Anthropic's context-editing docs (opens in a new tab), the trigger for clearing out old tool results defaults to 100,000 input tokens. So does the SDK's client-side compaction threshold. A million-token window shipped with a hundred-thousand-token default for throwing things away.
The length isn't what breaks it. The question is.
Long-context retrieval is close to solved. Ask a model to find one fact in a very long document and it will find it. What it can't reliably do is take three facts from three places in that document and put them together.
The cleanest demonstration holds the length constant and varies only the shape of the question. PredicateLongBench (opens in a new tab) runs its synthetic tasks at roughly 128K tokens. At high reasoning effort, Opus 4.6, GPT-5.4 and Gemini 3.1 score 87, 95 and 93 on a simple existence query. On global aggregation over that same context they score 23, 12 and 2. Add near-sorted decoys to a locate task and they score 1, 0 and 2. The paper notes Opus's misses on that first row were refusals rather than wrong answers, which probably understates it.
Nothing got longer. The question changed shape, and the floor gave way.
OpenAI built the GraphWalks benchmark for this exact gap, and its framing (opens in a new tab) is still the clearest statement of the problem I've found:
Few real-world tasks are as straightforward as retrieving a single, obvious needle answer. [Users need models to] retrieve and understand multiple pieces of information, and to understand those pieces in relation to each other.
Anthropic's numbers show the same gradient, and you can see it without leaving a single table.
| Model | Context | Parents (1 hop) | BFS (many hops) |
|---|---|---|---|
| Claude Opus 4.8 | 256K | 99.3 | 85.9 |
| Claude Opus 4.8 | 1M | 83.3 | 68.1 |
| GPT-5.5 | 256K | 90.1 | 73.7 |
| GPT-5.5 | 1M | 58.5 | 45.4 |
One hop is nearly solved. Several hops isn't, in either lab's model.
Chroma's Context Rot work (opens in a new tab), across 18 models, goes further. Degradation shows up on tasks with no retrieval difficulty at all, including plain text replication. Repeat this back to me, word for word, gets less reliable as the input grows. That closes the escape hatch people reach for, which is "my question is simple." Simple was tested.
It isn't dilution. It's interference.
Nobody has fully pinned the mechanism down. The best-supported candidate isn't attention spreading itself thin. It's competition between similar things.
One study across thirty models (opens in a new tab) found retrieval accuracy falling log-linearly with the number of competing items in context. In its regression, model size predicted resilience (p = 0.005) and context length didn't (p = 0.886). What the limit tracks is how much similar material the model has to sift, not how many tokens it was handed.
Chroma found something pointing the same way and genuinely strange: shuffled haystacks outperform logically structured ones, across all eighteen models. If this were dilution, coherent structure shouldn't hurt.
The finding that ties it to the rest of this post is that models don't maintain relational state as they read. They reconstruct it (opens in a new tab) when you ask. Probes for global state sit below 0.3 accuracy while query-local probes reach about 0.9. There's no running map of how things connect. There's a map rebuilt at query time, out of whatever is competing for attention at that moment.
Anthropic stopped publishing the curve
While checking the above I found something worth reporting on its own.
Anthropic's "what's new" page for Opus 5 (opens in a new tab) promises "consistent instruction following, tool calling, and reasoning throughout the window." The Opus 5 system card (opens in a new tab) publishes nothing you could check that against. Section 8.9 cites ProgramBench, an agentic coding benchmark. The words "MRCR" and "GraphWalks" appear zero times across 193 pages.
That would be unremarkable if the benchmark had been stable. It hasn't. The 4.6 card (opens in a new tab) reported MRCR and GraphWalks by context bin. The 4.8 card reports GraphWalks. The 5 card reports neither. The long-context benchmark changed twice across those three releases, which leaves no published Anthropic benchmark connecting Opus 5 back to Opus 4.6.
The 4.8 card is careful about this, and the carefulness is the point. It states plainly that it separates the 256K and 1M subsets unlike prior cards, which is a warning that its numbers and the earlier cards' numbers are not the same measurement. I've held to that rule here: every GraphWalks figure above comes from that one table.
And from page 200, a line I keep coming back to:
1M context subset results are not reproducible via the public API, as the problems exceed its 1M token limit.
Anthropic's million-token benchmark won't run through Anthropic's million-token API.
That same table is also the best evidence that this is improving. BFS at the 1M bin went 16.3 to 40.3 to 68.1 across Opus 4.6, 4.7 and 4.8. Multi-hop traversal at a million tokens roughly quadrupled in two releases. It just isn't being disclosed for the current flagship.
Every round number you've heard is a safety margin
You'll read that models are "reliable to about 100k." Treat that carefully, including the version of it I quoted above.
Anthropic's context-editing trigger defaults to 100,000. Coding agents compact somewhere below their window limit, and the published thresholds disagree with each other depending on whose reverse-engineering you're reading. None of those is a measured degradation point. They're margins chosen so an agent doesn't hit a wall mid-turn, and no vendor publishes degradation data justifying the figure.
The benchmarks that formally define an effective context length give numbers far below the folklore. NoLiMa (opens in a new tab) put GPT-4.1 at 16K and Claude 3.5 Sonnet at 4K. It hasn't been updated since June 2025 and never tested an Opus model.
Takeaway
So there's no published effective-context figure for any current frontier model. The round ones in circulation are engineering defaults being quoted as measurements, and that laundering is worth noticing whoever is doing it.
Relationships are what break, at both ends
Everything above is about relationships. The model finds the facts. It struggles to hold how they connect.
Now look at your own business.
The facts are already everywhere. Invoices, a website, three years of email, a calendar. What isn't written down anywhere is the connective tissue. Why you quote that type of customer high. Why the service page is built the way it is. Which jobs you've learned to walk away from, and the reasoning that got you there. What "sounds like us" means, precisely enough that somebody else could apply it.
Those are relationships too, and they live in one head. They're the class of thing that never makes it into a document, because people write down decisions and not the reasons behind them.
I want to be careful here, because it's the weakest joint in the argument. Graph edges in a token window and tacit business reasoning are not the same object, and I'm claiming a parallel, not an identity. The parallel is still the useful part: the thing models are worst at holding is the thing businesses are worst at recording. Both ends fail on structure rather than on facts.
Which is what a knowledge base is, and why "just paste your documents in" was never going to work. It isn't a pile of facts. The facts are the easy part and the model can already find them. It's the connections, made explicit, written down once, in a form small enough to hand over one task at a time.
The knowledge base exists. It's in the wrong place.
I maintain one for a painting contractor I've worked with since 2025. Several dozen files: a README that indexes the rest, dated records of decisions and how they were closed out, the history of what was asked for and what we agreed, written procedures for the jobs that keep recurring.
Nothing exotic. Markdown in version control, with a hook that stops anyone committing an unredacted database dump.
Almost none of it started as a document. It started as messages, calls and decisions that would otherwise have evaporated inside a week. Someone had to sit inside the relationship and decide, over and over, that a particular connection was worth a file. That judgment is the whole product. It's also why every knowledge base template you've downloaded is still sitting in your Downloads folder. A template can hold facts. It can't tell you which relationships matter.
Here's the uncomfortable half, and the reason I'm writing this at all. That knowledge base lives in my repository. The client writes his own site content with an LLM and, as far as I know, gives it none of this. What comes back reads like every other painting company's. The files exist. They've never met the tool that needs them.
Written down isn't sufficient. Written down, and in reach of the thing you're actually typing into. Two different jobs, and almost everybody stops after the first one, if they get that far.
The honest version of the counterargument
Start with the hole in the middle of it. Nobody has measured the quantity the argument turns on. There's no published number for how many relationships a model can hold, and no established protocol for measuring one. It's a shape rather than a figure, and I'm reasoning from a direction of travel.
The naive version is wrong, too. "The model runs out of relationship slots" is not what's happening, and anyone selling you a slot count is hand-waving.
The evidence is uneven, and a post that showed you only collapse curves would be cherry-picking. Some frontier models are close to length-insensitive on some benchmarks. People are bad at this as well: on one aggregation benchmark the human baseline was 75% at 128K and 50% at 1M. Some of what I've described is task difficulty, not model defect.
The strongest objection to the business half is a separate one, and it has nothing to do with models. Extraction may not stay manual. An adaptive interview plus ingestion of a company's site, invoices and email could plausibly draft a decent context pack with very little human editing. I haven't tested that properly, and I'm suspicious of my own incentive to disbelieve it, because I currently do this by hand and charge for it.
What I think happens next
I think the relational gap closes faster than most people expect. 16.3 to 68.1 across three releases is not a plateau, and if that continues, the model half of this argument has maybe two years left in it.
The business half has no such trajectory. No release note is going to write down why you quote that customer high.
So here's the prediction, held at the confidence a shape rather than a number deserves. Context windows stop being interesting. Retrieval turns into plumbing nobody thinks about. The gap that stays open is the boring one, and the businesses getting real value out of this won't be the ones with the best retrieval stack. They'll be the ones that wrote down how they work.
The tell is specific. Watch whether the next flagship system card ships a long-context degradation curve. If it does, the missing section was a blip and I read too much into it.