Someone on my team asked last month why we still had a retrieval pipeline running when "the model can just read the whole codebase now." Fair question. GPT-6 Astra takes 1.05M input tokens. Claude Opus 5.5 and Sonnet 5 take 1M. Gemini 3.8 Flash takes 1,048,576. On paper, most internal knowledge bases fit in a single prompt.
We kept the retrieval pipeline. Not out of habit: I checked the numbers, and for our case RAG was still cheaper and more accurate. But that answer doesn't generalize, and "it depends" isn't a framework, it's a shrug. This post is the actual framework: five variables that determine whether you should retrieve, stuff, or compact, with the tradeoffs made concrete instead of hand-waved.
If you haven't worked through the fundamentals of retrieval-augmented generation yet, start with the RAG lesson: I'm assuming you know what RAG is and building on top of that here, not re-explaining it. Same for context engineering, which covers the general practice of deciding what goes into a context window.
The three options, briefly
RAG: retrieve a small set of relevant chunks from an external store and put only those in context. Long context: put everything relevant in context and let the model find what it needs. Compaction: keep working in a live, growing context, but periodically compress older parts of it into a summary so the conversation can run longer than any single context window without a full re-retrieval step.
These aren't mutually exclusive, a production agent might do all three at once, retrieving documents, loading a few of them in full, and compacting its own scratchpad as it works. But most teams don't decide this deliberately. They default to whatever they built first and never revisit it. That's the actual problem worth fixing.
Five variables that actually decide this
1. Cost per query
This is the one people get wrong most often, because "context is basically free now" is true at the token-price level and false at the query-volume level.
Say you have a 500K-token knowledge base and you're on Claude Sonnet 5 at roughly $3 per million input tokens (check current numbers on the models page: they move). Stuffing the whole thing into every query costs you $1.50 in input tokens alone, before output. Run that 10,000 times a month and you're at $15,000/month just for input, before you've generated a single useful token of output. Prompt caching cuts this dramatically for repeated identical prefixes, but only if your knowledge base doesn't change between queries and your traffic pattern actually hits the same cache often enough to matter.
RAG flips the cost curve: you pay for embedding the corpus once (cheap, and mostly happens at ingest time) and then pay only for the 3-10 chunks (maybe 2-4K tokens) that are actually relevant to each query. At $3/M input tokens, that's fractions of a cent per query instead of $1.50.
The crossover point in practice: if your knowledge base is small enough to fit in context and stays roughly the same across most queries, long context with caching wins. If it's large, changes per query, or you're running high query volume against a stable subset of a bigger corpus, RAG wins on cost by one or two orders of magnitude. Do the arithmetic for your actual token count and query volume before deciding: don't reason about it in the abstract.
2. Data freshness
RAG's retrieval step queries a live index, so it can serve documents updated seconds ago. Long context means whatever you loaded into the prompt is frozen the moment you built it: if the underlying data changes mid-conversation, the model doesn't know.
For a support bot answering questions against documentation that updates weekly, that's not a big deal: reload the context occasionally. For a system answering questions about live inventory, account status, or anything that changes within the span of a conversation, stuffing a static snapshot into context is actively wrong, not just inefficient. That's a retrieval problem by definition, not a context-size problem, no amount of context window growth fixes staleness, because staleness is about when data was fetched, not how much of it you fetched.
3. Traceability and citations
This is the most underrated axis and the one I've seen bite teams hardest in production.
RAG gives you a paper trail for free: each answer traces back to the specific chunks that were retrieved, so you can show "this claim came from document X, section Y" and a human can verify it. That matters enormously for anything with compliance, legal, or medical exposure, and it matters practically for debugging: when the model gets something wrong, you can inspect exactly what it saw.
Long context doesn't give you this automatically. The model read a million tokens and produced an answer; unless you explicitly instruct it to cite sources and structure the document with IDs it can reference (more on this in the companion post on prompting large contexts), you have no idea which part of the input actually drove the output. You can engineer citations into a long-context setup, but it's work you get for free with RAG's retrieval step.
If your use case needs an audit trail ("show your work" for a regulator, a customer, or your own postmortem process) that alone can be reason enough to keep retrieval in the loop even when everything would technically fit in context.
4. Latency
Bigger prompts take longer to process, even before generation starts. Time-to-first-token scales with input size, and on a 500K-token prompt that processing overhead is real, often multiple seconds before the model produces anything, even with prompt caching absorbing repeat costs on cached prefixes. RAG's retrieval step also takes time (a vector search is typically tens of milliseconds), but the resulting prompt is small, so the model's own processing is fast.
For anything synchronous and user-facing (a chat interface, a live agent loop) the latency difference between "search then generate on 3K tokens" and "generate on 500K tokens" is often the whole difference between a responsive product and one that feels broken. For batch or async workloads, this matters much less.
5. Context window economics over a session
This is where compaction earns its place, and it's a genuinely different problem from the RAG-vs-stuffing question above. Even a 1M-token window fills up over a long enough session, a multi-hour coding agent run, a long customer support conversation, a research task that accumulates tool outputs turn after turn. Eventually you hit the ceiling regardless of how big the window is.
Claude's compaction (currently in beta) solves this specific problem: rather than you writing custom summarization logic, the API can automatically summarize older turns once input tokens cross a threshold: the default trigger is 150,000 tokens, with a 50,000-token minimum if you want it to kick in earlier. There are two flavors: on-demand compaction, where your application decides when to request a summary, and threshold compaction, where the API triggers it mid-request once you cross your configured limit. Both are beta features on the current model line (Fable 5.1, Opus 5.5, Sonnet 5, and a few prior versions): check the compaction docs for the current compatibility list before building on it, since beta feature support shifts between model releases.
Compaction is not a substitute for RAG or for careful context curation. It's an answer to a different question: not "what should be in context for this query" but "how do I keep a long-running session from ever hitting a hard wall." Use it alongside RAG or long-context stuffing, not instead of either.
The decision framework, put together
Ask these in order:
- Does the data change faster than your context gets rebuilt? If yes, you need retrieval or a live tool call: long context alone can't fix staleness.
- Do you need per-claim citations or an audit trail? If yes, lean RAG, or engineer explicit citation structure into your long-context prompt (see the companion post).
- What's the query volume against this data? High volume against a corpus larger than a few hundred K tokens: RAG wins on cost by a wide margin. Low volume against a stable, moderate-size corpus, long context with caching is simpler and often cheaper in practice, since you skip building and maintaining a retrieval pipeline.
- Is this synchronous and user-facing? If latency is visible to a human waiting on the response, favor the smaller effective context RAG gives you.
- Is this a long-running session rather than a one-shot query? If so, add compaction to whichever of the above you chose: it's orthogonal, not competing.
Most real systems end up as a combination: RAG for the corpus-scale knowledge, long context for the handful of documents directly relevant to the current task, and compaction to keep the session itself from running out of room. The mistake isn't picking the "wrong" one of these, it's treating it as a single binary choice and never revisiting it as your query volume, data freshness needs, or session length change.
One more thing worth doing before you commit to an architecture: actually measure it. Build a small golden test set of representative queries and check retrieval quality against long-context quality on your own data, not on a benchmark, see building evaluation datasets for how to set that up. The "it depends" answer is unsatisfying because it's usually followed by nobody actually testing which one is true for their case.
For the execution side of this, how to structure a long-context prompt so it actually performs well once you've decided to use one: see prompting with a million-token context window. And for picking which model to route a given query to once you've settled on an architecture, see LLM routing.



