I loaded a 400K-token codebase into Claude last month, asked it to find every place a deprecated config flag was still referenced, and it missed three of them: all sitting in the middle third of the file list. Not because the model can't handle 400K tokens. Because I dumped the files in alphabetical order with no structure and expected the model to treat token 200,000 with the same attention as token 5,000. It doesn't, and pretending otherwise is the most common mistake I see people make with these context windows.
This post assumes you've already decided to put a large amount of content in context, if you're still weighing that against retrieval, read the companion post on RAG vs long context vs compaction first. This one is about execution: given that you're loading a huge context, how do you structure it so the model actually uses it well.
The context windows you're actually working with
As of this writing: GPT-6 (Astra, Sol, Luna) takes 1.05M input tokens with a 128K max output. Claude's current line (Fable 5.1, Opus 5.5, Opus 5, Sonnet 5) takes 1M input tokens, also 128K max output (Haiku 4.5 is smaller, at 200K input, with no separate large-output tier). Gemini 3.8 Flash and 3.5 Flash-Lite take 1,048,576 input tokens with 65,536 max output. Full current numbers, updated as they change, live on the models page and the model comparison table.
A million tokens is roughly 750,000 words: a shelf of novels. The size stopped being the bottleneck a while back. What the model does with that size is the actual variable now.
Structure and position: what the vendors actually recommend
This isn't a hunch, Anthropic's own prompting docs are explicit about it: place long documents near the top of the prompt, above your instructions and query, not buried after them or interleaved with other content. Their internal testing found response quality improved measurably (up to 30% in some evaluations) just from moving the query to the end after the documents instead of before them. That's a free win that costs you nothing but reordering.
For multiple documents, wrap each one in explicit tags with a source identifier:
<document source="q3-earnings-call-transcript.pdf">
[content of document 1]
</document>
<document source="q3-board-deck.pdf">
[content of document 2]
</document>
<document source="analyst-notes-jane-doe.txt">
[content of document 3]
</document>
This does two things. First, it stops the model from blending content across documents when you ask it to compare or attribute something, without clear boundaries, a model working across a dozen documents will sometimes synthesize a claim that's actually a merge of two separate sources. Second, it gives you a citation anchor: ask the model to reference <source> when it makes a claim, and you get a rough audit trail even without a retrieval pipeline. That's the citation problem from the RAG comparison post, you can partially engineer your way around it in a long-context setup by just being disciplined about document tagging.
Label everything. If you're loading code, include file paths as comments. If you're loading transcripts, timestamp and speaker-tag them. The model isn't reading your context the way you'd read a book start to finish, it's doing something closer to pattern-matching across the whole thing at once, and unlabeled content gives it nothing to anchor a pattern-match to.
The lost-in-the-middle problem
This is real, well-documented, and not something later model generations have fully solved. The original finding (from a 2023 Stanford/Berkeley/Samaya AI paper, since replicated in various forms) is that models reliably do better at using information placed at the very beginning or very end of a long context than information buried in the middle. It shows up across tasks: retrieval, multi-document QA, ranking, arithmetic reasoning.
The leading explanation isn't that the model "forgets" the middle in some literal sense, it's that beginning and end positions act as strong signals for what the task actually is, while the middle gets treated more like undifferentiated background. Chroma's 2026 research on this (they call it "context rot") extends the finding: it's not just position, it's that performance degrades progressively as total context length grows, even on tasks that should be position-independent, and even well under the model's stated context limit.
Practical mitigations that actually help:
- Put the most important content at the start or end, never buried in the middle of a large stack of documents. If you have one document that matters most, don't file it fourth in a stack of ten.
- Restate the critical instruction after the content, not just before it. If your query is 50 tokens and your context is 400K tokens, that instruction is easy for the model's attention to under-weight relative to everything else. Repeating the core ask right before generation (after all the documents) measurably helps.
- Chunk and label rather than wall-of-text. A context window structured as fifteen clearly delimited documents outperforms the same content pasted as one undifferentiated blob, because the model can use the boundaries as retrieval anchors.
- For agentic loops, compact aggressively. If you're accumulating tool outputs and conversation turns over a long session, the oldest, least relevant material is exactly what tends to end up "in the middle" by the time you're deep into the session. This is a second, complementary reason (beyond hitting a hard token ceiling) to use something like Claude's compaction: it keeps the active context shorter and more front-loaded with what actually matters right now.
When a huge context window doesn't help
Two failure modes show up even when everything is technically well within the window's stated limit.
Single-needle vs multi-needle retrieval. Needle-in-a-haystack benchmarks that ask a model to find one fact in a huge context tend to score very well across current frontier models: often near-perfect at 1M tokens. But real workloads are rarely single-needle. Ask a model to synthesize five related facts scattered across a large context, and accuracy drops substantially compared to the single-fact case: evaluations putting the gap at anywhere from 15 to 40 points aren't unusual. If your actual task is "find and combine several pieces of information," don't extrapolate confidence from a single-needle benchmark number: test your specific multi-fact retrieval pattern directly.
The cost of reasoning tokens on huge contexts. If you're using a model with extended or adjustable reasoning effort (see prompting reasoning models for the fundamentals), be aware that reasoning tokens are billed as output tokens even though they're hidden from the response, and a model reasoning over a genuinely large context tends to generate more of them, not fewer, because there's more material to work through before it commits to an answer. "Just load everything and let it think it through" can quietly turn into a very expensive query once you account for reasoning overhead on top of the input token cost. Test actual cost per query on your workload before assuming bigger context is a free upgrade.
Context rot on tasks well under the limit. As mentioned above, degradation isn't a cliff at the token ceiling: it's a gradual slope that starts well before you hit the window's stated maximum. A task that would perform reliably at 50K tokens of well-curated context doesn't necessarily perform just as reliably at 400K tokens of loosely curated context, even though both are comfortably inside a 1M window. Bigger available context is not the same claim as "load more and it'll do at least as well." Test at the size you're actually planning to use, not just under the theoretical ceiling.
A practical checklist
- Documents at the top, instructions and query at the end, restated once after the content
- Every document explicitly tagged with a source identifier
- Most important material at the start or end of the document stack, never buried in the middle
- For agentic/multi-turn use, compact old turns rather than letting context grow unbounded: see the compaction docs for current setup
- If the task involves combining multiple facts, test multi-needle performance directly rather than trusting single-needle benchmark numbers
- Budget for reasoning-token overhead on large contexts if you're using an effort-adjustable model: check prompting reasoning models
- Before defaulting to "just load everything," run the cost and latency math from the RAG vs long context framework: a bigger window doesn't mean stuffing it is the right call for every case
The size of these windows genuinely changed what's practical: you can now do things that needed a retrieval pipeline eighteen months ago. But size alone doesn't get you good results. Structure does.



