llm-integration.eu

Claude Usage Limits: Plans vs API vs Bedrock 2026

Claude usage limits 2026: five-hour plan sessions, API Start/Build/Scale caps of 500 to 200,000 USD, plus Bedrock, Google and Foundry quotas for EU teams.

Updated 12 min readFacts verified on 10 October 2026

TL;DR

Claude usage limits are three different systems. Chat plans use a five-hour session pool plus a weekly cap. The API uses Start, Build and Scale tiers with spend caps of 500, 1,000 or 200,000 USD. Bedrock, Google Cloud and Foundry each keep their own quotas. For EU processing, use Bedrock eu.anthropic.claude-sonnet-5-5, not the Claude API.

What are Claude usage limits on chat plans?

Claude usage limits on claude.ai are a conversation budget, not a request-per-minute quota. Paid plans have a session limit that resets every five hours and a weekly limit that resets on a fixed weekday. Chat, Claude Code and Claude Desktop share the same pool. There is no published message count. Length limits are the context window, a different constraint.

The usage and length article, read on 10 October 2026, is explicit. Usage is how much you can work across all chats before you wait. Length is how much one chat can hold. Paid newest models support up to a 1M token context window. Others support 500K or 200K. A slice of that window is reserved for the reply.

The pricing FAQ states the multipliers. Free covers everyday questions. Pro gives at least 5x more usage per 5-hour session than Free. Max 5x and Max 20x give 5x or 20x more than Pro per session. Team Standard seats give more than Pro. Premium seats give 5x more than Standard, which matches the 1.25x and 6.25x Pro figures in the Team help article. Limits are per member. One person at the cap does not freeze the rest of the team.

Plan Session multiplier Weekly cap Extra after the cap
Free baseline not a paid weekly pool wait or upgrade
Pro (20 USD month, 17 USD annual) at least 5x Free yes, all models usage credits at API rates
Max 5x / 20x (100 / 200 USD) 5x or 20x Pro yes, plus a separate Fable weekly cap usage credits; 100 or 200 USD monthly API credits
Team Standard (25 / 20 USD) 1.25x Pro yes, per seat org-level usage credits
Team Premium (125 / 100 USD) 6.25x Pro yes, per seat; Fable is 50% of the weekly pool org-level usage credits
Enterprise (20 USD seat, annual) none on usage-based seats none every token at API rates

Sources: claude.com/pricing, Pro, Max, Team, Enterprise. Seat prices are US list prices and exclude tax. Team needs 2 to 150 people. Self-serve Enterprise starts at 20 seats, sales-assisted at 50.

Usage-based Enterprise is the odd one out. The Enterprise article says the seat fee covers access only. There are no per-seat usage limits and no included token allowance. Admins set spend limits instead. Older Enterprise orgs that still show Standard and Premium seats keep those caps until renewal.

Usage credits, described in the credits article, switch a paid plan to pay-as-you-go at standard API rates once a session or weekly cap is gone. They are not the monthly API credits bundled with Max and Team. Those API credits do not pay for Claude, Claude Code or Cowork after the plan pool is empty. Seat prices and the full ladder sit in Claude pricing plans 2026. Team admin detail is in the Claude Team plan for Europe.

What are Claude API rate limits and spend caps?

Claude API rate limits are organization-level RPM, ITPM and OTPM per model class, plus a monthly spend cap per usage tier. Start stops at 500 USD, Build at 1,000 USD, Scale at 200,000 USD. Custom has no published cap. A 429 on the spend cap has no retry-after until 00:00 UTC on the first of the next month.

The rate limits page from this run lists three standard tiers. New or low-history orgs may start in Evaluation, below these numbers, and move up automatically. Limits use a token bucket. A 60 RPM ceiling can fire as 1 request per second, so a burst still 429s. Rate limits are shared across inference_geo: "us" and "global".

Tier Monthly spend cap Sonnet 5.5 / Opus 5.5 RPM ITPM OTPM Fable 5.x RPM / ITPM / OTPM
Start 500 USD 1,000 2,000,000 400,000 1,000 / 500,000 / 100,000
Build 1,000 USD 5,000 5,000,000 1,000,000 2,000 / 1,500,000 / 300,000
Scale 200,000 USD 10,000 10,000,000 2,000,000 4,000 / 4,000,000 / 800,000

Sonnet 5, Haiku 5.5 and Haiku 4.5 share the same Start, Build and Scale numbers as Sonnet 5.5 and Opus 5.5. Cache reads do not count toward ITPM on most models. Cache writes and uncached input do. The docs’ example: with 2,000,000 ITPM and an 80% cache hit rate you can process 10,000,000 total input tokens in a minute. Haiku 3.5 still counts cache reads. max_tokens does not count toward OTPM on the first-party API.

Message Batches have a separate queue. Start: 1,000 RPM, 200,000 requests in the processing queue, 100,000 per batch. Build: 2,000 / 300,000 / 100,000. Scale: 4,000 / 500,000 / 100,000. Managed Agents sit aside from Messages: 300 create RPM, 1,200 read RPM.

A spend cap you set yourself below the tier returns HTTP 400, not 429. Workspace overrides can sit below the org values. They cannot sit above. You cannot put overrides on the Default Workspace. Workspace work, keys and the Admin API live in the Claude Console guide. Token list prices live in Claude API pricing in the EU.

Read the configured limits instead of hardcoding them:

curl "https://api.anthropic.com/v1/organizations/rate_limits" \
  -H "x-api-key: $ANTHROPIC_ADMIN_KEY" \
  -H "anthropic-version: 2023-06-01"

That call is the Rate Limits API. It needs an Admin API key or an org:admin token. Workspace-scoped keys fail. Use it at gateway start and on a schedule.

Handle a 429 that still has headroom (rate limit, not spend cap):

# retry-after is seconds. The spend-cap 429 has no such header.
retry=$(curl -sS -D - -o /tmp/claude-body \
  https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{"model":"claude-sonnet-5-5","max_tokens":256,"messages":[{"role":"user","content":"ping"}]}' \
  | awk 'tolower($1)=="retry-after:" {print $2; exit}')
if [ -n "$retry" ]; then sleep "$retry"; fi

We recommend this split. Put evals and playground traffic in the Console. Put production with personal data on Bedrock or Google Cloud. The first-party API has no EU inference_geo.

How do Bedrock, Google Cloud and Foundry quotas differ?

Cloud quotas are not the Claude API tiers. Bedrock counts tokens per model and Region, with a burndown on bedrock-runtime. Google publishes QPM and separate input and output TPM per endpoint. Foundry publishes RPM, ITPM and OTPM per Azure subscription type. None of those numbers move when you change a Claude Console tier.

Amazon Bedrock keeps two endpoints. On bedrock-runtime, input and output share one TPM quota. Output tokens burn faster than input. The token burndown page from this run: Claude 4.8 output burns at 15x; Claude Opus 5.5, Sonnet 5, Opus 5 and Fable 5.1 burn at 10x; Claude 4.7 and earlier burn at 5x. Cache reads do not count. At request start Bedrock deducts input plus max_tokens. A high max_tokens can 429 you before the model writes a word. bedrock-mantle uses separate input and output TPM, so burndown does not apply there.

The Sonnet 5.5 model card (launch 28 September 2026, 1M context, 128k max output) publishes the EU geo ID eu.anthropic.claude-sonnet-5-5. From Frankfurt the destination Regions are Frankfurt, Stockholm, Milan, Spain, Ireland and Paris. London as source also keeps London. Default TPM values sit in Service Quotas and can sit below the published defaults. Account age and payment history change them.

Google’s Claude quota page says ML processing stays in the EU on Europe regional and multi-region endpoints. Sonnet 5.5 on a multi-region endpoint: 1,250 QPM, 12,500,000 uncached-plus-cache-write input TPM, 1,250,000 output TPM, 1,000,000 context. The global endpoint doubles those figures (2,500 / 25,000,000 / 2,500,000). Opus 5.5 multi-region is 1,000 / 10,000,000 / 1,000,000. Models launched after 26 May 2026 share a lineage bucket (anthropic-claude-sonnet, anthropic-claude-opus). Global and eu buckets are independent.

Foundry’s Claude quota page is the tightest default we read. Pay-as-you-go Sonnet 5.5 and Opus 5.5: 40 RPM, 40,000 ITPM, 8,000 OTPM. Enterprise and MCA-E: 10,000 / 10,000,000 / 2,000,000. Free Trial is 0 / 0 / 0. Fable on pay-as-you-go is also 0. Data Zone Standard exists for the US, not for the EU. Cache reads do not count toward ITPM.

Platform What you hit Sonnet 5.5 default (this run) EU processing?
Claude API Start RPM + ITPM + OTPM + 500 USD cap 1,000 / 2M / 400k No. inference_geo is us or global
Bedrock eu. TPM (input+output) + burndown on runtime Service Quotas, account-specific Yes. eu.anthropic.claude-sonnet-5-5
Google Europe multi-region QPM + input TPM + output TPM 1,250 / 12.5M / 1.25M Yes on eu
Foundry pay-as-you-go RPM + ITPM + OTPM 40 / 40k / 8k No EU data zone

Setup and first call: Amazon Bedrock in the EU, Claude on Vertex AI in Europe, Azure AI Foundry in Europe. GDPR routes: Claude GDPR comparison.

What happens when you hit a limit, and how do you raise it?

A chat-plan cap pauses that person until the five-hour or weekly reset, unless usage credits are on. An API rate limit returns 429 with retry-after. An API spend cap returns 429 with enforced_spend_limit_reached and no retry header until next month. Cloud 429s are quotas. The fix is backoff, then a quota ticket, not a Claude plan upgrade.

Do this in order:

  1. Read the error. Chat: Settings, Usage shows the next reset. API: look for retry-after versus enforced_spend_limit_reached. Bedrock and Foundry: HTTP 429, then CloudWatch or the Foundry Quota page.
  2. On a paid chat plan, enable usage credits under Settings, Usage, prepay a balance, and set a monthly spend cap. Mobile-store subscribers must do this on the web app.
  3. On the API, open Settings, Rate limits and use Request tier increase. Claude Platform on AWS has no that button. Mail Anthropic support with peak ITPM, OTPM and cache share.
  4. On Bedrock, raise Cross-Region InvokeModel tokens per minute for the model in Service Quotas. AWS says it will then offer the on-demand TPM and tokens-per-day pair. Priority goes to accounts that already consume their allocation.
  5. On Google Cloud, filter Quotas by anthropic-claude-sonnet or anthropic-claude-opus and request the eu multi-region metrics, not the global ones, if residency matters.
  6. On Foundry, open the Quota page and submit the increase form. Approval is not guaranteed. Free Trial stays at zero until you leave that offer.
  7. Cut wasted quota: prompt cache (reads skip ITPM on most models), Batch at 50% token price, and a realistic max_tokens on Bedrock runtime.

Sonnet 5.5 list price on the Anthropic pricing page is 2 USD input and 10 USD output per million tokens. Opus 5.5 is 4 / 20. Haiku 5.5 is 0.10 / 0.50 up to 100,000 prompt tokens and 0.50 / 2.50 above that. US-only inference (inference_geo: "us") on Claude 4.6 and later is 1.1x. Regional and multi-region cloud endpoints from Sonnet 4.5 add 10%. Batch halves both sides. Fast mode on Opus 5.5 is 8 / 40 and exists only on the first-party API.

A 200k cached prefix plus a 50-token question counts as 50 uncached input tokens toward API ITPM. That is why caching is a limit tool, not only a cost tool.

Which limit system should an EU team pick?

Pick the system that matches the data, not the marketing name. Chat plans are for people. The API is for products that may leave the EEA. Bedrock EU or Google Europe is for products that must stay in member states. Foundry is a US or global quota pool with a very low pay-as-you-go default.

We recommend this default for a European company:

  • Staff chat and light Claude Code: Team Standard, Premium for the few people who hit the weekly cap. Turn on org usage credits with a hard monthly number.
  • Heavy coding with no EU-processing mandate: Max 20x for one person is cheaper than a Team Premium seat if you only have one power user. Two or more people belong on Team.
  • Production APIs with personal data: Bedrock eu.anthropic.claude-sonnet-5-5 or Google eu multi-region. Request quota before go-live. Do not wait for the first 429.
  • Production APIs without a residency mandate: Claude API Scale, or Foundry Enterprise if Azure already holds the contract. Skip Foundry pay-as-you-go for anything above a demo. 40 RPM will not survive a launch.

The data residency page still offers only "us" and "global" for inference_geo. Workspace geo is "us" only and cannot change after create. Rate limits are shared across geos. That is why raising a Console tier does not create an EU region.

Claude Platform on AWS starts on the Start tier and bills in Claude Consumption Units at 0.01 USD each (100 CCU = 1.00 USD). Per-workspace rate limits and fast mode are not available there. Foundry bills the same CCU unit.

FAQ

How much do Claude usage limits cost to raise?

Chat plans do not sell a bigger session as a line item. You buy a higher plan or you buy usage credits at API rates. The API raise is a tier move: Start 500 USD, Build 1,000 USD, Scale 200,000 USD calendar-month caps, then Custom with sales. Bedrock, Google and Foundry quota tickets are free to file. You still pay the tokens you run.

Claude usage limits vs API rate limits: what is the difference?

Claude usage limits on a plan are a five-hour plus weekly conversation budget shared by chat and Claude Code. API rate limits are RPM, ITPM and OTPM plus a monthly spend cap on an organization. A Team Premium seat at 6.25x Pro does not raise your Console Start tier. A Scale org at 10,000 RPM does not give staff more claude.ai messages.

Do Claude usage limits apply in the EU the same way?

Yes for the numbers. No for residency. Plan pools and API tiers are global. inference_geo is still only us or global. Workspace geo is us. EU processing is a different product: Bedrock eu. profiles or Google Europe endpoints, with their own quotas.

What happens if one Team member hits a Claude usage limit?

The other seats keep working. That person waits for the five-hour or weekly reset, or the org enables usage credits. The reset weekday is fixed per account and shown under Settings, Usage. It does not move when someone starts using Claude.

Does prompt caching raise effective Claude API limits?

Yes on most models. Cache reads do not count toward ITPM. Cache writes and uncached input do. The rate-limits docs walk through 2,000,000 ITPM with an 80% hit rate as 10,000,000 total input tokens per minute. Haiku 3.5 is the exception that still counts reads. Caching also cuts the bill: Sonnet 5.5 cache hits are 0.10 USD per million tokens.

Why did Bedrock 429 after I raised max_tokens?

On bedrock-runtime, the first deduction is input tokens plus max_tokens. A 32,000 max_tokens on a short answer reserves that many tokens up front. Set max_tokens near the real completion size. After the response, Bedrock adjusts to input plus output times the burndown (10x on Sonnet 5 and Opus 5.5).

Sources

  1. Anthropic docs: Rate limits (10 October 2026)
  2. Claude Help Center: How do usage and length limits work? (10 October 2026)
  3. Claude: Pricing (10 October 2026)
  4. Claude Help Center: What is the Team plan? (10 October 2026)
  5. Claude Help Center: What is the Enterprise plan? (10 October 2026)
  6. Anthropic docs: Pricing (10 October 2026)
  7. Anthropic docs: Data residency (10 October 2026)
  8. AWS: How tokens are counted in Amazon Bedrock (10 October 2026)
  9. AWS: Claude Sonnet 5.5 model card (10 October 2026)
  10. Google Cloud: Quotas for Anthropic Claude models (10 October 2026)
  11. Microsoft Learn: Claude model quotas and rate limits (10 October 2026)
  12. Anthropic docs: Rate Limits API (10 October 2026)

Related guides