If you're building on the Claude API, rate limits are the first operational constraint you hit. The error messages are clear enough, but the underlying tier system — why you're limited, what moves you to the next tier, and how to work within the limits productively — isn't well-documented in one place.
Here's what you need to know.
How Anthropic's rate limit tiers work
Anthropic structures access in five tiers. You start at Tier 1 when you first add a credit card. Limits increase automatically as you spend more and maintain your account in good standing.
As of mid-2026, the approximate tier structure (verify current limits at console.anthropic.com — these change):
| Tier | Spend threshold | RPM (Sonnet 4.6) | Input TPM |
|---|---|---|---|
| Tier 1 | First payment | 50 | 40,000 |
| Tier 2 | ~$100 spent | 1,000 | 80,000 |
| Tier 3 | ~$500 spent | 2,000 | 160,000 |
| Tier 4 | ~$5,000 spent | 4,000 | 400,000 |
| Tier 5 | Enterprise contract | Custom | Custom |
RPM = requests per minute. TPM = tokens per minute (input side).
Important: limits vary by model. Haiku 4.5 has higher default limits than Sonnet 4.6, which has higher limits than Opus 4.8 and Fable 5. If you're hitting limits on Sonnet 4.6, check whether Haiku handles the task acceptably — you get higher throughput at a fraction of the cost.
Understanding the three limit types
When you hit a rate limit, you get a 429 Too Many Requests response. The response headers tell you which limit was hit:
x-ratelimit-limit-requests: your RPM ceilingx-ratelimit-limit-tokens: your TPM ceilingx-ratelimit-remaining-requests: requests left in this windowx-ratelimit-remaining-tokens: tokens left in this window
The three limit types you'll encounter:
Rate per minute (RPM): too many requests in a 60-second window. Common when parallelizing many short requests.
Tokens per minute (TPM): too many input + output tokens in a 60-second window. Common when processing long documents or using large context windows.
Tokens per day (TPD): daily token budget exhausted. Primarily a Tier 1 issue — the daily cap is tight on the lowest tier.
Retry patterns
Don't retry immediately on a 429. Use exponential backoff:
import anthropic
import time
client = anthropic.Anthropic()
def call_with_retry(messages, max_retries=5):
for attempt in range(max_retries):
try:
response = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=1024,
messages=messages
)
return response
except anthropic.RateLimitError:
if attempt == max_retries - 1:
raise
wait_time = (2 ** attempt) + 1 # 2, 3, 5, 9, 17 seconds
time.sleep(wait_time)
The retry-after header in the 429 response gives you the minimum wait time. Use it as your floor. The Anthropic Python SDK also has built-in retry logic you can configure:
client = anthropic.Anthropic(
max_retries=5,
)
Strategy 1: Batch with the Messages Batches API
For non-real-time workloads, use the Batches API instead of individual API calls. Batch requests are processed asynchronously, aren't subject to the same synchronous rate limits, and cost 50% less per token.
batch = client.messages.batches.create(
requests=[
{
"custom_id": f"req-{i}",
"params": {
"model": "claude-sonnet-4-6",
"max_tokens": 1024,
"messages": [{"role": "user", "content": text}]
}
}
for i, text in enumerate(documents)
]
)
# Poll for results
while True:
result = client.messages.batches.retrieve(batch.id)
if result.processing_status == "ended":
break
time.sleep(60)
If you're processing documents, running evaluations, or doing any task where latency isn't critical, batching is the correct approach. You stay under rate limits and pay half the price.
Strategy 2: Prompt caching for repeated context
If your system prompt is the same across many requests (which it usually is), enable prompt caching. Cached tokens cost ~10% of normal input prices and don't put the same load on TPM limits.
response = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=1024,
system=[{
"type": "text",
"text": your_long_system_prompt,
"cache_control": {"type": "ephemeral"}
}],
messages=[{"role": "user", "content": user_message}]
)
The cache TTL is 5 minutes for ephemeral caches. For a system prompt that's the same across requests, caching it brings the effective cost of Sonnet 4.6 much closer to Haiku 4.5 pricing.
Strategy 3: Route by task complexity
Not every task needs the same model. A classification task that takes one token of reasoning doesn't need Sonnet 4.6. Haiku 4.5 handles:
- Classification and routing decisions
- Simple extraction (pull field X from document Y)
- Format conversion
- Short summarization
Save Sonnet 4.6 for tasks requiring multi-step reasoning, code generation, and nuanced synthesis. Save Fable 5 for tasks requiring maximum capability.
This "model routing" approach lets you run significantly higher throughput at the same cost, because Haiku's limits are higher than Sonnet's.
Strategy 4: Chunk large documents
TPM limits bite hardest on large documents. Instead of sending a 100k-token document in one request, chunk it:
def chunk_document(text, chunk_size=8000, overlap=500):
words = text.split()
chunks = []
for i in range(0, len(words), chunk_size - overlap):
chunk = " ".join(words[i:i + chunk_size])
chunks.append(chunk)
return chunks
Process chunks sequentially or in controlled parallel batches to stay within TPM limits. For most extraction tasks, 8k-token chunks work well.
Managing TPM vs RPM
TPM is usually the harder limit to manage. A single request with a 100k-token document uses more TPM budget than 100 simple chat messages.
Key TPM management tactics:
- Trim conversation history: don't include more context than needed for the current request
- Use
max_tokenssensibly: set it to the maximum you expect, not the maximum allowed - Streaming doesn't change your TPM, but it reduces perceived latency — sometimes allowing you to spread requests more naturally over time
Requesting a limit increase
Tier increases happen automatically as you spend, but you can contact Anthropic to request an early increase for legitimate use cases. Email enterprise@anthropic.com or use the console to request higher limits. Include:
- Your use case and expected volume
- Why you need limits beyond your current tier
- Timeline for when you'll need them
For enterprise scale (beyond Tier 4), the self-service tiers cap out. Anthropic's enterprise contracts offer custom rate limits, SLAs, and dedicated capacity. Start the procurement process before you need it — it takes time.
For related cost optimization techniques, see our prompt caching guide and Anthropic Batch API guide.



