llm-integration.eu

Claude Agent SDK in the EU: Bedrock and Vertex AI

Run the Claude Agent SDK on Amazon Bedrock or Vertex AI with EU data residency: env vars, eu. model IDs, hosting patterns, cost drivers and Python/TS code.

Updated 12 min readFacts verified on 1 October 2026

TL;DR

The Claude Agent SDK runs Claude Code’s agent loop inside your own Python or TypeScript service. For EU data residency, skip the Anthropic API: set CLAUDE_CODE_USE_BEDROCK=1 with an eu. inference profile from Frankfurt, or CLAUDE_CODE_USE_VERTEX=1 with CLOUD_ML_REGION=eu. Pin the model, isolate the container, and budget for tokens, not compute.

What is the Claude Agent SDK, and why does the provider matter?

The Claude Agent SDK is a library that embeds Claude Code’s tools, agent loop, permissions, sessions and hooks into an application you operate. The model provider is a separate choice. By default the SDK calls api.anthropic.com. Two environment variables switch it to Amazon Bedrock or Google Cloud, where you choose an EU routing option.

The Agent SDK overview describes it as “Claude Code as a library”. It ships for Python (claude-agent-sdk) and TypeScript (@anthropic-ai/claude-agent-sdk). Prerequisites are Python 3.10 or newer, or Node.js 18 or newer. Both packages bundle a native Claude Code binary, so the CLI version moves with the SDK version. At the time of writing the Python changelog tops out at 0.2.163 with bundled CLI 2.1.286, and the TypeScript changelog at 0.3.287, at parity with Claude Code v2.1.287.

The architecture decides most of what follows. Each call to query() spawns a claude subprocess that talks to your code over stdio and owns a shell, a working directory and JSONL transcripts on local disk. That is not a stateless API wrapper. It is a long-lived process with files, which is why hosting and data protection need more thought than for a plain Messages API call.

Three things are worth knowing before you pick a provider:

  • No subscription login. Anthropic does not allow third-party products built on the SDK to offer claude.ai login or plan rate limits unless previously approved. Production agents authenticate with an API key or cloud credentials.
  • Managed Agents is not an option on the clouds. Anthropic’s hosted alternative, Claude Managed Agents, is listed as not supported on Bedrock and on Google Cloud. If your residency strategy rests on Bedrock or Vertex, you host the SDK yourself.
  • Same agent, different endpoint. The SDK documentation states that outbound HTTPS goes to api.anthropic.com or to your provider’s regional endpoint when you run on Bedrock or Google Cloud. Your agent code does not change.

We recommend Bedrock or Vertex over the direct API whenever agents touch personal data or source code that falls under a GDPR processing record, mainly because the processing contract, the region control and the bill then sit in a cloud account you already govern. For the general provider comparison, see our Claude on AWS Bedrock in the EU guide.

How do you point the Agent SDK at Bedrock in the EU?

Set CLAUDE_CODE_USE_BEDROCK=1, give the process AWS credentials and set AWS_REGION to an EU region such as eu-central-1. Claude Code maps any eu-* region to the eu. cross-region inference profile prefix. Then pin the model explicitly, because the unpinned default changes with SDK releases and currently resolves to Opus 5.5.

The Bedrock setup page documents the resolution order for the region: AWS_REGION, then AWS_DEFAULT_REGION, then the active AWS profile, then a fallback to us-east-1. That last step matters. A container without a region variable and without an AWS config file ends up in the US. Credentials come from the standard AWS SDK chain: an instance or task role, an SSO profile, access keys, or a Bedrock API key in AWS_BEARER_TOKEN_BEDROCK.

These are the EU profile IDs you would pin today:

Model Bedrock EU profile What AWS says Single-Region option
Claude Sonnet 5 eu.anthropic.claude-sonnet-5 Keeps data within EU regions bedrock-mantle in eu-north-1 or eu-west-1
Claude Opus 5.5 eu.anthropic.claude-opus-5-5 Keeps data within EU regions No EU region on bedrock-mantle
Either, global. prefix global.anthropic.claude-... Routes worldwide, no residency constraint Not applicable

The Opus 5.5 model card lists a launch date of 22 September 2026, a 1M token context window and 128K max output. Both models list Frankfurt as a source Region for the EU geo profile. AWS also notes that cross-Region profiles can route to opt-in Regions you never enabled, where prompts and outputs may be stored for abuse detection. For an EU profile that is still inside the EU, but your processing record should say “EU regions”, not “Frankfurt”.

A minimal Python agent on Bedrock, read-only and bounded:

import asyncio
from claude_agent_sdk import query, ClaudeAgentOptions, ResultMessage

async def main():
    options = ClaudeAgentOptions(
        cwd="/work/session-a",
        allowed_tools=["Read", "Glob", "Grep"],
        permission_mode="dontAsk",
        max_turns=20,
        env={
            "CLAUDE_CODE_USE_BEDROCK": "1",
            "AWS_REGION": "eu-central-1",
            "ANTHROPIC_MODEL": "eu.anthropic.claude-sonnet-5",
        },
    )
    async for message in query(prompt="Summarize the files in this directory", options=options):
        if isinstance(message, ResultMessage):
            print(message.subtype, message.total_cost_usd)

asyncio.run(main())

In Python, env is merged on top of the inherited environment, so the instance role stays visible. Setting ANTHROPIC_MODEL also moves background tasks onto that model. The IAM policy needs bedrock:InvokeModel, bedrock:InvokeModelWithResponseStream, bedrock:ListInferenceProfiles and bedrock:GetInferenceProfile. One functional gap: the WebSearch tool is not available on Bedrock. If your agent needs web research, give it an MCP tool behind your own egress proxy. The same settings apply to Claude Code on developer laptops, covered in Claude Code on Bedrock in the EU.

How do you run it on Vertex AI with the EU multi-region?

On Google Cloud set CLAUDE_CODE_USE_VERTEX=1, ANTHROPIC_VERTEX_PROJECT_ID and CLOUD_ML_REGION=eu. The value eu selects the multi-region host aiplatform.eu.rep.googleapis.com. Do not use a single region such as europe-west1 for current models: Anthropic states that specific regional endpoints support Sonnet 4.6 and earlier only.

Google’s model pages for Claude Opus 5.5 and Claude Sonnet 5 both list the Europe multi-region for model availability and for ML processing. The default quota on the multi-region is lower than on the global endpoint. For Opus 5.5 it is 1,000 queries per minute, 10,000,000 input tokens per minute and 1,000,000 output tokens per minute, half of the global endpoint’s figures. Sonnet 5 gets 1,250 QPM and 12,500,000 input TPM on the multi-region. An agent that fans out into parallel subagents can hit those limits, and the SDK hosting docs explicitly warn that wide subagent fan-outs can run into rate limits.

The fallback is the trap here too. If CLOUD_ML_REGION is unset or malformed, Claude Code falls back to us-east5. Set it explicitly in the container spec, not in someone’s shell profile.

The same agent in TypeScript on Vertex:

import { query } from "@anthropic-ai/claude-agent-sdk";

for await (const message of query({
  prompt: "Summarize the files in this directory",
  options: {
    cwd: "/work/session-a",
    allowedTools: ["Read", "Glob", "Grep"],
    permissionMode: "dontAsk",
    maxTurns: 20,
    env: {
      ...process.env,
      CLAUDE_CODE_USE_VERTEX: "1",
      CLOUD_ML_REGION: "eu",
      ANTHROPIC_VERTEX_PROJECT_ID: "your-project-id",
      ANTHROPIC_MODEL: "claude-sonnet-5",
    },
  },
})) {
  if (message.type === "result") console.log(message.subtype, message.total_cost_usd);
}

In TypeScript env replaces the subprocess environment, so spread process.env or you lose PATH and the Google credentials. The service account needs roles/aiplatform.user. Anthropic’s feature list for Google Cloud includes the server-side web search tool, which Bedrock lacks. Test whether your SDK version exposes it before you rely on it. The residency details of Claude on Google Cloud are in our Claude on Vertex AI Europe article.

How should you host an agent in production?

Treat every agent session as a process with state, not as a request. One session is one subprocess, its transcripts sit on local disk, and none of that survives a restart unless you mirror it to durable storage. Pick a session pattern first, then harden the container, then decide how tenants are separated.

The hosting guide names four patterns:

Pattern Container lifetime Fits State requirement
Ephemeral One container per task Invoice extraction, bug fix, translation None beyond output files
Long-running Persistent, many sessions Email triage, chat bots Size RAM for peak concurrent sessions
Hybrid Spins down when idle, resumes Support agents, multi-day research SessionStore is mandatory
Multi-agent Several SDK processes in one container Simulations, collaborating agents Separate cwd and settings per agent

The documented starting point is 1 GiB RAM, 5 GiB disk and 1 CPU per agent, and Anthropic calls that a floor. Memory grows with session length, so measure peak RSS for a realistic session and size hosts with “agents per host = (host RAM minus overhead) / per-session ceiling”.

For EU-sensitive workloads we would apply five controls:

  1. Run in a private subnet with egress only to an internal proxy that allowlists the Bedrock runtime host in your region or aiplatform.eu.rep.googleapis.com. The secure deployment guide shows a container with --network none, --cap-drop ALL, --read-only and a mounted proxy socket.
  2. Persist transcripts in an EU bucket through a SessionStore adapter, and alert on mirror_error messages, because a failed mirror write drops that batch and continues.
  3. Separate tenants: setting_sources=[], CLAUDE_CODE_DISABLE_AUTO_MEMORY=1, a per-tenant CLAUDE_CONFIG_DIR and a per-tenant cwd. Without this, one customer’s CLAUDE.md context can leak into another’s session.
  4. Bound every session with max_turns. The SDK has no top-level session timeout.
  5. Export telemetry with CLAUDE_CODE_ENABLE_TELEMETRY=1 and the OTLP variables to a collector you host. Prompt text and tool inputs are excluded from exports by default.

On permissions, dontAsk plus an explicit allowed_tools list is the headless setting we would start from. Do not combine allowed_tools with bypassPermissions and assume the list restricts anything: the permissions page warns that bypass mode approves unlisted tools too, including Bash.

What drives the cost of an EU agent?

Tokens. Anthropic’s own hosting guidance says token cost typically dominates container cost by an order of magnitude or more: a minimal container runs about 0.05 USD per hour, while one long session can spend dollars. The EU route adds a 10% premium over global endpoints on both Bedrock and Google Cloud for current models.

Anthropic list prices from the pricing page, with the documented 10% regional premium applied. AWS and Google publish the binding prices on their own pages, so check those before you budget.

Model Input per MTok (list) Output per MTok (list) With 10% EU premium
Claude Sonnet 5 2 USD 10 USD 2.20 / 11 USD
Claude Opus 5.5 4 USD 20 USD 4.40 / 22 USD
Claude Haiku 4.5 1 USD 5 USD 1.10 / 5.50 USD

Our worked example, not a vendor figure: an agent session that consumes 500,000 uncached input tokens and 50,000 output tokens costs about 1.65 USD on Sonnet 5 via an EU endpoint and about 3.30 USD on Opus 5.5. A full day of minimal container time is roughly 1.20 USD. Per session, the model choice moves the bill more than the hosting choice.

Four cost drivers to control:

  • The unpinned default. On Bedrock and Vertex, Claude Code’s default primary model is Opus 5.5. The docs warn that unpinned deployments are billed at the Opus rate. Pin Sonnet 5 unless the task needs Opus.
  • Cache expiry. On Bedrock and Google Cloud the prompt cache TTL defaults to 5 minutes. Agents that run many short sessions with longer gaps pay full input price each time. ENABLE_PROMPT_CACHING_1H=1 trades higher write cost for more cache hits.
  • Subagent fan-out. total_cost_usd includes subagents, usage does not. Read cost from total_cost_usd or model_usage.
  • Missing caps. max_budget_usd (TypeScript maxBudgetUsd) stops a call when its own spend crosses the limit.

total_cost_usd is a client-side estimate from a price table bundled with the SDK. It does not know your AWS or Google discount, and the docs say not to bill users from it. Use AWS Cost Explorer or Google Cloud billing for the real number. For a cross-provider token comparison see our Claude API pricing EU overview.

Run it yourself or outsource?

Writing the agent is the small part. Running it 24/7 is the larger cost: someone owns the container platform, the IAM boundary, the egress proxy, SDK upgrades, model pins and on-call. If you do not already run containers with that discipline, an operator can make sense. If you do, keep it in-house.

What in-house operation involves for this topic, based on the documented failure modes:

Recurring task Why it recurs Who typically owns it
SDK and CLI upgrades Bundled binary is pinned to the SDK version; the docs advise taking patches continuously and reading the changelog before a minor Platform engineer
Model pin reviews Defaults moved from Opus 5 to Opus 5.5 in v2.1.280; new EU profiles appear per model Platform engineer plus product owner
Permission boundary reviews Tool allowlists, deny rules and hooks drift as agents get new jobs Security
Monitoring and on-call No session timeout, memory growth, mirror_error, quota 429s Operations
IAM and egress policy Region pinning, inference profile ARNs, proxy allowlists Cloud team

Our estimate, not a sourced figure: for a single production agent with customer data, plan for a fraction of a platform engineer’s time on an ongoing basis plus a share of an on-call rotation. That is cheap if you already have both. It is expensive if you have to build them for one agent.

Outsourcing makes sense when the agent runs around the clock on personal data, you have no on-call, and no one owns AWS or Google Cloud IAM. It makes less sense for a CI job or an internal batch agent that runs a few times a day: an ephemeral container with no persistent state has little to operate.

If you outsource, demand these points in writing:

  1. SLA on the agent runtime, separate from the model provider’s SLA. The provider controls containers and session storage, AWS or Google controls the model endpoint.
  2. DPA (AVV) as processor, covering transcripts, working directories and telemetry, not only prompts.
  3. Subprocessor list naming the cloud, the exact region or multi-region, and any sandbox or observability vendor.
  4. Access model: who can read transcripts and session storage, how break-glass access is logged, and whether the Bedrock or Vertex account is yours or theirs. We would insist on your own cloud account.
  5. Exit: transcripts in your bucket, infrastructure as code handed over, no proprietary session format. If the agent runs in your account, exit means revoking a role.

FAQ

Does the Claude Agent SDK work with Amazon Bedrock?

Yes. Set CLAUDE_CODE_USE_BEDROCK=1, provide AWS credentials and set AWS_REGION. With an eu-* region, Claude Code prefers eu. inference profiles such as eu.anthropic.claude-sonnet-5. Pin the model with ANTHROPIC_MODEL so the default does not shift to Opus 5.5 unnoticed. WebSearch is not available on Bedrock.

What does an agent session cost in the EU?

It depends on tokens, not containers. At list price plus the 10% regional premium, Sonnet 5 costs 2.20 USD input and 11 USD output per million tokens. Our example session with 500,000 input and 50,000 output tokens lands at about 1.65 USD. A minimal container costs roughly 0.05 USD per hour.

Claude Agent SDK vs Managed Agents: which fits EU data?

The Agent SDK, if your residency plan relies on Bedrock or Google Cloud. Managed Agents runs the loop on Anthropic’s infrastructure and is listed as not supported on either cloud. With the SDK you host the loop in your own EU account and pick the EU endpoint yourself.

Which Google Cloud region should I use for Claude agents?

CLOUD_ML_REGION=eu, the Europe multi-region. Opus 5.5 and Sonnet 5 are available there. Specific regions such as europe-west1 serve Sonnet 4.6 and earlier only. An unset region falls back to us-east5.

Can my agent use a Claude Pro or Max login?

No. Anthropic does not allow third-party products built on the Agent SDK to offer claude.ai login or plan rate limits without prior approval. Use an API key or Bedrock or Google Cloud credentials.

Is total_cost_usd what I will be billed?

No. It is a client-side estimate from a price table bundled at build time. It does not reflect cloud discounts or every billing rule. Use AWS Cost Explorer or Google Cloud billing for invoices, and keep total_cost_usd for budgets and alerts.

Sources

  1. Claude Code docs: Agent SDK overview (1 October 2026)
  2. Claude Code docs: Agent SDK quickstart (1 October 2026)
  3. Claude Code docs: Hosting the Agent SDK (1 October 2026)
  4. Claude Code docs: Track cost and usage (1 October 2026)
  5. Claude Code docs: Agent SDK permissions (1 October 2026)
  6. Claude Code docs: Securely deploying AI agents (1 October 2026)
  7. Claude Code docs: Amazon Bedrock (1 October 2026)
  8. Claude Code docs: Google Cloud Agent Platform (1 October 2026)
  9. AWS: Claude Sonnet 5 model card (1 October 2026)
  10. AWS: Claude Opus 5.5 model card (1 October 2026)
  11. AWS: Supported Regions and models for inference profiles (1 October 2026)
  12. Google Cloud: Claude Opus 5.5 on Agent Platform (1 October 2026)
  13. Google Cloud: Claude Sonnet 5 on Agent Platform (1 October 2026)
  14. Anthropic: Claude on Google Cloud (1 October 2026)
  15. Anthropic: Pricing (1 October 2026)
  16. claude-agent-sdk-python CHANGELOG (1 October 2026)
  17. claude-agent-sdk-typescript CHANGELOG (1 October 2026)

Related guides