I've watched three different teams get surprised by their AI bill this year, and it was never because the per-token price was wrong. It was because nobody on the team could explain, in advance, why a "cheap" model ended up costing more than the expensive one. Per-token pricing is the easy part: it's printed on a page. What's hard is knowing which number on that page actually applies to your workload, and that's the part vendors don't put in a table.
This isn't another price list. The site already has one, current model pricing and specs and the interactive cost-at-scale comparison are both kept up to date and will always beat a static table in a blog post that goes stale. What I want to walk through here is the reasoning: how to go from "$2 per million input tokens" to "what will this actually cost me a month," and the three places that estimate usually goes wrong: reasoning tokens, effort settings, and caching.
The token price you see isn't the token price you pay
Every major vendor prices in dollars per million tokens, split into input and output, and output is always priced higher: often 5x the input rate. As of this check (September 23, 2026), the spread across current models looks like this:
- OpenAI: GPT-6 Luna at $0.10 / $0.50 per million, GPT-6 Sol at $2 / $10, GPT-6 Astra at $10 / $50
- Anthropic: Claude Haiku 4.5 at $1 / $5, Claude Sonnet 5 at $2 / $10, Claude Opus 5 at $5 / $25, Claude Opus 5.5 at $4 / $20, Claude Fable 5.1 at $10 / $50
- Google: Gemini 3.5 Flash-Lite at $0.30 / $2.50, Gemini 3.8 Flash at $0.75 / $3.75 (introductory, through December 31, 2026: rising to $1.50 / $7.50 after)
For the full current table with context windows and max output, see the models page, I'm not going to re-derive it here because it'll be out of date the moment a vendor ships an update, which happens more often than you'd think. Opus 5.5 landed the day before I'm writing this, at a lower price than Opus 5. That's not unusual.
The output multiplier is the first thing people underprice in their heads. If your workload generates long responses, a coding agent writing files, a research summary, a JSON extraction with a big schema: output tokens dominate the bill even when your prompts are short. A system that sends 200 input tokens and gets back 2,000 output tokens is spending 10x more on the response than the request, before you've touched any other lever.
Reasoning tokens are the hidden multiplier
Here's the part that actually changes people's bills: reasoning (or "thinking") tokens are billed as output tokens, even though you often never see them in the response.
All three vendors now expose some form of effort or thinking control:
- OpenAI's
reasoning.effortparameter (none through max) - Anthropic's
output_config.effort(low / medium / high / xhigh / max), paired withthinking: {type: "adaptive"} - Google's
thinking_level(minimal / low / medium / high)
These controls exist because reasoning models generate an internal scratchpad before producing the final answer, and that scratchpad costs real output-token money. Turn effort up on a "cheap" model and you can multiply its effective cost several times over, sometimes past what a pricier model would have cost at a lower effort setting, because the expensive model needed less reasoning to get the same answer right.
This is the single most common way teams get surprised. They benchmark a model at default settings, ship it, then someone bumps reasoning.effort to high for a tricky edge case and the average request cost doubles or triples without anyone noticing until the invoice arrives. If you're tracking cost per request, effort level needs to be part of that tracking: not just model name and token count.
The fix isn't "always use low effort." It's matching effort to task. Simple classification, extraction, and formatting tasks rarely benefit from high effort: you're paying for reasoning the model doesn't need. Multi-step agentic work, hard coding problems, and anything where a wrong answer is expensive to unwind justify the higher setting. I go deeper on when reasoning actually helps versus when it's wasted spend in prompting reasoning models: worth reading before you set a default effort level for a whole product.
Caching pays off faster than most people assume
The second lever, and the one with the best cost-to-effort ratio, is prompt caching. If your requests share a stable prefix, a system prompt, a set of tool definitions, a long document you're asking multiple questions about, caching that prefix means you pay full price once and a fraction of that price on every subsequent read.
All three vendors support this, and the discount is steep enough that it's worth restructuring your requests around it. As of this check, cached reads run at roughly a 90% discount off the base input rate on both Anthropic's and Google's APIs, the exact multiplier and minimum cacheable prefix length vary by model and vendor, so treat that as an order-of-magnitude figure and check current docs before you build cost projections around it. Writing to the cache typically costs a small premium over the standard input rate (Anthropic prices this at roughly 1.25x), which is why caching only pays off once you're reading the same prefix more than a couple of times.
The catch is that caching is a strict prefix match. Any change earlier in the request, a timestamp in the system prompt, a reordered tool list, a non-deterministic field, invalidates everything after it, and you silently fall back to full price without any error telling you so. If you're not sure your caching is actually working, check the response's cache-read token count directly rather than assuming it is. I've seen production systems run for months with caching quietly broken because someone put datetime.now() in a system prompt.
Where caching earns its keep: a coding agent that re-sends the same large system prompt and tool schema on every turn, a support bot with a long static knowledge document, a batch job asking many questions about one big input. Where it doesn't help: one-off requests, or workloads where every prompt is genuinely different from the last.
Batch APIs: the other free lunch, if you can tolerate the delay
If your workload doesn't need a response in real time, nightly data processing, bulk classification, generating a large set of summaries, every major vendor offers a batch API at roughly half the standard price, in exchange for accepting an asynchronous turnaround (typically up to 24 hours rather than immediate). Batch and caching stack, so a workload that's both cacheable and batchable can see combined savings well past what either lever delivers alone.
The tradeoff is operational, not just financial: batch APIs mean building a submit-and-poll workflow instead of a synchronous request, and your results can come back out of order, so you need to key them by request ID rather than position. For anything user-facing this doesn't work. For backend jobs it's close to free money sitting on the table.
Worked example: estimating your own workload
Say you're building a support-ticket classifier that runs on 50,000 tickets a month. Each ticket averages 300 input tokens (the ticket text plus a system prompt) and the model returns a 20-token category label: no reasoning needed, this is pure classification.
Without caching: 50,000 × 300 = 15M input tokens, 50,000 × 20 = 1M output tokens. On a fast-tier model like GPT-6 Luna ($0.10/$0.50) that's $1.50 input + $0.50 output = $2 a month. Trivial.
Now say instead you're running an agentic coding assistant that handles 500 sessions a month, each session averaging 40,000 input tokens (accumulated conversation + file context, much of it a stable system prompt and repeated file reads) and 8,000 output tokens, at high reasoning effort on a flagship model like Claude Opus 5 ($5/$25). Without caching: 500 × 40,000 = 20M input tokens = $100, plus 500 × 8,000 = 4M output tokens = $100. That's $200/month before reasoning inflation, and at high effort on hard problems, output token counts commonly run 2-4x higher than the "final answer" length alone, because the reasoning tokens are counted too. Realistically you're looking at $300-500/month for that workload, and that's exactly the kind of number that surprises a team that budgeted off the sticker price alone.
Now add caching. If 30,000 of those 40,000 input tokens per session are a stable system prompt and tool schema that doesn't change turn to turn, and you're re-reading it on every one of, say, 10 turns per session, moving that stable portion behind a cache breakpoint turns most of those reads into $0.50-per-million-equivalent tokens instead of $5-per-million tokens. On a workload with that much repeated prefix, caching alone can cut the input-token bill by more than half.
The lesson isn't "here's the formula", it's that your actual monthly cost depends on three things the sticker price doesn't show you: how much of your input is genuinely new each request versus cacheable, how much reasoning your task actually needs versus what effort setting you've defaulted to, and whether your traffic is bursty enough to benefit from batching. Model the workload, not the token price.
Where to go from here
For the exact current numbers, use the models page and the cost-at-scale calculator, the calculator specifically lets you plug in your own volume assumptions rather than eyeballing a per-token rate. If you're trying to decide which model tier fits a given task in the first place (before you even get to cost), the decision guide walks through that by input type, task type, and budget. And if reasoning effort is the lever you haven't tuned yet, prompting reasoning models covers when it's worth the spend and when it's just burning tokens on a problem the model didn't need to think hard about.
One caveat on everything above: prices in this space move fast, and vendors change them with a blog post and no warning. I checked every number here against current vendor documentation on September 23, 2026 (Anthropic's docs, OpenAI's developer platform, and Google's Gemini API docs) but by the time you're reading this, something may have shifted. Cache-discount and batch-discount magnitudes especially are the kind of number vendors adjust without much fanfare, so verify against current docs before you build a cost model you're going to bet a budget on.



