Your assistant worked fine for two weeks. Then it started ignoring a rule that's plainly in the system prompt, quoting a policy that was updated in March, and answering slower. Nothing in the prompt changed. What changed is that the context quietly filled up with history, retrieved chunks, tool output and a few things nobody remembers adding.
A context window budget is a written plan for what's allowed into the window, how many tokens each part may use, and how you decide to cut. The decision rule that makes it work is simple: a block earns its tokens only if removing it makes your test cases worse. Everything below is how to apply that rule without guessing.
This is a method post. For the concept, read what is context engineering. For code that counts tokens and trims history, see token counting and context management. Here you get the audit, the worksheet and a small ablation script.
Why a bigger window doesn't remove the need for a budget
Two findings are worth knowing, and both have primary sources.
Anthropic's engineering team describes context as a finite resource with diminishing returns, and uses the term context rot for recall getting worse as the window fills. Their goal is stated as the smallest set of high-signal tokens that makes the outcome likely (Effective context engineering for AI agents).
Chroma's Context Rot study tested 18 models and found performance dropped as input grew even on simple tasks. Two details matter for budgeting. Distractors that are topically close to the answer hurt more as inputs get longer. And on a long-memory benchmark, a focused prompt of roughly 300 tokens beat the full roughly 113k-token version for every model family they tried.
Position matters too. The Lost in the Middle paper found models do best when the relevant information is at the start or end of the input and worse when it's buried in the middle.
None of this says "never use long context". It says relevance beats volume, and you only find out what's relevant by testing.
Step 1: What's your working ceiling and your reserve?
Write down two numbers before you add anything.
Reserve for output. If the model has to write a 1,500-word report, or think before answering, that space is spoken for. Reserve the max output setting plus room for reasoning, not the average response.
Working ceiling. Pick a total input size you'll aim to stay under, well below the model's maximum. There's no correct number. Start somewhere you're comfortable (say a quarter of the advertised window), run your test cases at that size, and move it only if you have evidence. Raising the ceiling should require a reason. Lowering it requires none.
Step 2: How do you inventory everything in the window?
Print the actual assembled prompt for a real request, not the template. Then list every block. Most systems have the same seven candidates:
- System prompt (role, rules, tone, format)
- Tool definitions
- Retrieved documents or chunks
- Conversation history
- User or account data
- Few-shot examples
- Tool results from earlier steps
You'll often find an eighth: something nobody can explain. A stale paragraph pasted in during an incident, a duplicate of the policy that also arrives via retrieval, a 40-line tone guide. Write each block on its own row with its approximate token count. Four characters per token is a fine rough estimate for English prose; use your provider's counter when you need exact figures.
Step 3: Which blocks earn their tokens?
Ask three questions per block, in this order.
Does it change the answer? Run the ablation in the next section. This is the only question that needs data.
Does it need to be here every time? Plenty of material is true but only sometimes relevant: the returns policy matters on returns questions. Anything that applies to a minority of requests should be fetched on demand (retrieval, or a tool the model calls), not carried permanently. Anthropic's guidance describes the same split: keep lightweight references and load detail at runtime, with a hybrid of upfront and on-demand often working best.
Is it the right size? A raw account record with 60 fields where the task uses 6 is the classic leak. Select fields before they reach the window. The same goes for tool results: cap them, and say in the result that it was capped.
A rule of thumb for what usually survives: instructions the model would otherwise get wrong, facts it can't know, and two or three diverse examples. What usually gets cut: full history past the last few turns, duplicated policy text, rules written to fix an incident two model versions ago, and tool definitions for tools this task never calls.
How do you test what a block is worth? A 30-line ablation
Ablation means removing one thing at a time and measuring. You need 10 to 20 test cases with an expected answer or a checkable property. If you don't have these, build them first; golden test sets covers how, and promptfoo can run the comparisons.
Below is the loop. I ran it with a fake "model" that only answers if the needed fact appears in the context, so it checks the logic of the script, not any real model. You supply call_model and passes.
def approx_tokens(text: str) -> int:
return max(1, round(len(text) / 4)) # rough: ~4 chars per token for English
def run_ablation(blocks, cases, call_model, passes):
"""blocks: dict name -> text. cases: list of (question, expected).
call_model(context, question) -> answer. passes(answer, expected) -> bool."""
def score(ctx_blocks):
context = "\n\n".join(ctx_blocks.values())
return sum(passes(call_model(context, q), e) for q, e in cases)
baseline = score(blocks)
rows = []
for name in blocks:
reduced = {k: v for k, v in blocks.items() if k != name}
rows.append((name, approx_tokens(blocks[name]), score(reduced) - baseline))
return baseline, rows
With the toy model, it printed this:
baseline passes: 2/2
block ~tokens delta if removed
old_history 555 0 CUT
tone_guide 230 0 CUT
policy 36 -1 KEEP
pricing 22 -1 KEEP
Read the output carefully, because one line is a trap. tone_guide shows "no change" only because my test cases check facts, not tone. If tone matters, add cases that check tone (a rubric or a judge), otherwise the ablation will happily tell you to delete it. Ablation measures what your tests measure, nothing more.
Two more cautions. Run it with at least a handful of runs per case if your model is non-deterministic, since a one-case swing can be noise. And ablate blocks together when they overlap: if the policy arrives both in the system prompt and via retrieval, removing either alone shows no change, but removing both breaks things.
The worksheet
Copy this table, fill one row per block, and keep it next to the prompt in your repo.
| Block | Source | Tokens now | Needed on every request? | Score change if removed | Cap / rule | Decision |
|---|---|---|---|---|---|---|
| System prompt | file | yes | max N, one owner | keep / trim | ||
| Tool definitions | code | only for tools used | load by task | keep / split | ||
| Retrieved chunks | search | no | top K, min relevance | keep / tighten | ||
| History | session | partly | last N turns + summary | keep / summarize | ||
| Account data | database | partly | fields X, Y, Z only | trim | ||
| Examples | file | yes | 3 diverse | keep | ||
| Tool results | runtime | no | cap N tokens | cap |
Below the table, write four lines: the output reserve, the working ceiling, the total of the "tokens now" column, and the date you last ran the ablation. If the total is above the ceiling, something has to go, and the "score change" column tells you what.
A worked example (illustrative numbers)
Take a support assistant for a SaaS product. These figures are made up to show the arithmetic, using the rough four-characters-per-token estimate. They aren't measurements from a real deployment.
| Block | Before | After | What changed |
|---|---|---|---|
| System prompt | 4,200 | 1,800 | Removed duplicated policy and a stale incident rule; kept rules the ablation showed mattered |
| Tool definitions | 6,000 | 2,500 | 14 tools down to the 6 this assistant ever calls |
| Retrieved chunks | 6,000 (10 chunks) | 2,400 (4 chunks) | Top 4 above a relevance threshold |
| History | 14,000 | 3,000 | Last 6 turns verbatim plus a short summary |
| Account record | 6,000 | 1,200 | Six fields instead of the full record |
| Examples | 0 | 750 | Added 3 short examples; outputs got more consistent |
| Tool results | 12,000 | 5,000 | Capped, with a "truncated" note |
| Total | 48,200 | 16,650 | About two-thirds smaller |
Notice the "after" column adds something. Budgeting isn't only cutting. If an ablation shows a missing block would help (here, examples), it goes in the plan with a cap like everything else.
In what order should the remaining context go?
Ordering is the second half of the budget, and three facts decide it.
Long documents go first, instructions and the question last. Anthropic's prompting guidance recommends placing long inputs near the top and the query at the end, and reports that ending with the query can improve quality by up to 30 percent in their tests on complex multi-document inputs (prompting best practices). Treat that as a thing to test on your data, not a guarantee.
Stable material goes before changing material. Prompt caching matches an identical prefix in the order tools, system, then messages, and a change to anything earlier invalidates the cache from that point on (prompt caching docs). A timestamp near the top of the system prompt can defeat caching on every call. Our prompt caching guide walks through the cost side.
The rule that must not be lost goes at an edge. Given the lost-in-the-middle result, put the constraint that must hold at the start, and restate it close to the question if breaking it is expensive. Don't put your most important retrieved chunk fifth of ten.
When should you re-run the audit?
Budgets rot like documentation does. Re-run the ablation when you change model, when you add a tool, when retrieval content changes, and whenever a session-length graph shows input tokens creeping upward. Put a date on the worksheet. A budget that's six months old describes a system you no longer run.
For agents that run long, trimming at assembly time isn't enough; you also need compaction and notes that survive a reset. That's covered in Claude context compaction for long-running agents and the context engineering lesson.



