llm-integration.eu

Agentic RAG with Claude in the EU 2026

Agentic RAG with Claude in the EU: Bedrock Managed Knowledge Bases in Frankfurt or Ireland, AgenticRetrieveStream pricing, and when to host the index yourself.

Updated 10 min readFacts verified on 8 October 2026

TL;DR

Agentic RAG with Claude in the EU means the index lives in Frankfurt or Ireland, the retriever plans sub-queries, and the answer cites sources. Use a Managed Knowledge Base and AgenticRetrieveStream. The Anthropic API has no EU inference. Send tokens through the EU geo profile eu.anthropic.claude-sonnet-5.

How is agentic RAG different from classic RAG?

Classic RAG retrieves once and stuffs chunks into the prompt. Agentic RAG splits the question, searches more than once, checks whether the hits are enough, and pulls the full document when it must. On Bedrock that loop is AgenticRetrieveStream. It runs only on managed knowledge bases, not on a vector store you operate yourself.

AWS describes six steps: load session history, plan sub-queries, retrieve against up to five retrievers, call GetDocumentContent when a full document is needed, synthesise an answer, return results and citations. A speculative retrieve often runs on the raw question before the first plan, so the first round is not empty. generateResponse defaults to true. Set it to false if your own orchestrator will hand the chunks to Claude.

That is a different API from RetrieveAndGenerate. That call retrieves once and generates with citations. AWS states it cannot be used with managed knowledge bases. On the managed index you get Retrieve (one hybrid search) or AgenticRetrieveStream (several iterations). maxAgentIteration has a minimum of 2. Fewer iterations cut cost and can stop a complex question early.

AWS’s own example splits “Which magazine was started first, Arthur’s Magazine or First for Women?” into two founding-date queries. An internal policy question about a DPA, retention and deletion works the same way. One vector hit rarely covers it.

The Anthropic API can run the same loop. A tool returns search_result blocks and Claude cites them when citations.enabled is true. Search results are part of the standard Messages API, no beta header, on every active model except Claude Haiku 3. Citations are off by default. The EU catch: inference_geo accepts only us and global. Workspace geo is us only. If personal data must stay in the EU, call Claude on Bedrock or Google Cloud, not api.anthropic.com.

Which Bedrock Knowledge Base runs in the EU?

A managed knowledge base exists in Europe in Frankfurt (eu-central-1), Ireland (eu-west-1) and London (eu-west-2). For residency inside the Union pick Frankfurt or Ireland. London is the United Kingdom. AWS runs ingestion, storage, embeddings and rerank. You supply the sources.

The region list names eight Regions worldwide, including those three in Europe. Knowledge Bases with structured data stores (natural language to SQL) also exist in Zurich and Paris. That is a different product: tables, not unstructured PDFs.

AWS contrasts managed and customer-managed knowledge bases. Agentic retrieval, seven native connectors and the service-managed embedder sit only on the managed side. Customer-managed lets you pick the vector store, but it has no agentic retrieval and only S3 plus Custom as sources.

Feature Managed Knowledge Base Customer-managed
Agentic retrieval yes no
Connectors 7 (S3, SharePoint, Confluence, Web Crawler, Google Drive, OneDrive, Custom) S3 and Custom
Embeddings and rerank service-managed, no extra charge your own Bedrock embedding
Vector store AWS operates it you operate it
Query API Retrieve or AgenticRetrieveStream Retrieve or RetrieveAndGenerate
EU Regions (unstructured) Frankfurt, Ireland, London depends on the store you pick

Use Sonnet 5 as the generator on bedrock-runtime with the EU geo profile eu.anthropic.claude-sonnet-5. The model card lists Knowledge Bases as supported there and unsupported on bedrock-mantle. Mantle is the single-Region path in Ireland or Stockholm. Agentic RAG and single-Region inference do not overlap on Sonnet 5. EU geo from Frankfurt routes to Frankfurt, Stockholm, Milan, Spain, Ireland and Paris. That stays in the EU and is not a pin to one Region. AWS warns that cross-Region inference can share data across Regions. Region tables and IAM live in our Amazon Bedrock EU guide.

Embeddings if you choose them: Titan Text Embeddings V2 (amazon.titan-embed-text-v2:0) in eu-central-1, eu-west-1, eu-west-2, eu-west-3 and more, at 256, 512 or 1024 dimensions. Titan G1 stays at 1536 dimensions and fewer EU Regions. Cohere Embed English and Multilingual sit in Frankfurt, Ireland, London and Paris at a fixed 1024 dimensions. The Data Automation parser for images is Oregon-only and in preview. For EU RAG stay on a text parser or a foundation-model parser that exists in the source Region.

What does agentic RAG cost on Bedrock?

Index storage is 5.00 USD per GB of raw data per month. A standard search is 1.00 USD per 1,000 Retrieve calls. Agentic retrieval is 4.00 USD per 1,000 AgenticRetrieveStream calls plus 1.00 USD per 1,000 underlying Retrieve calls. Service-managed embeddings and the service-managed reranker are listed at 0 USD on the pricing page.

AWS works two examples on 50 GB and about 100,000 documents. 100,000 standard searches: 250 USD storage plus 100 USD retrieve, 350 USD a month. 100,000 agentic calls that each make two underlying Retrieve calls: 250 plus 400 plus 200, 850 USD a month. Parser and rerank are included in both examples. Claude tokens for the visible answer are extra once you pick your own foundation model. Type MANAGED folds planning into the agentic price. Type CUSTOM adds the provider’s model rates.

Line item Price
Index storage 5.00 USD per GB of raw data / month
Standard Retrieve 1.00 USD per 1,000 calls
Agentic Retrieve (managed LLM) 4.00 USD per 1,000 calls plus 1.00 USD per 1,000 underlying Retrieve
Agentic Retrieve (your LLM) 1.00 USD per 1,000 underlying Retrieve plus model price
Managed embeddings and rerank 0 USD
Bedrock Data Automation as parser 0.010 USD per page of standard output
Claude Sonnet 5.5 (API) 2 USD input / 10 USD output per 1M tokens
Regional premium from Sonnet 4.5 10 percent over global

Anthropic pricing lists Sonnet 5.5 at 2 / 10 USD and Opus 5.5 at 4 / 20 USD per million tokens. From Sonnet 4.5, Haiku 4.5 and Opus 4.5, regional and multi-region endpoints on Bedrock and Google Cloud add 10 percent over global. That is the cost of the EU profile. Batch halves both sides. A cache hit on Sonnet 5.5 is 0.10 USD per million, 0.05x. US-only inference_geo on the Anthropic API is 1.1x and does not solve residency.

Cohere Rerank 3.5, if you replace the service-managed reranker, is 2.00 USD per 1,000 queries. One query may hold up to 100 document chunks. 350 chunks count as 4 queries. Each chunk is at most 512 tokens including the query. Generator token rates sit in our Claude API pricing comparison.

Default quotas from the quota page: 10,000 managed knowledge bases per account per Region (adjustable), 200 data sources (fixed), 50 concurrent ingestion jobs (fixed), 10 TB of raw data per knowledge base (fixed), 10,000 characters per query, 600 Retrieve RPM per knowledge base with a burst of 25 RPS, 300 AgenticRetrieveStream RPM per account. 10 TB at 5 USD/GB would be 51,200 USD in storage alone. Most teams stay far below that.

How do you start AgenticRetrieveStream in Frankfurt?

Create the knowledge base in eu-central-1, attach S3 or SharePoint, wait for ingestion, then call AgenticRetrieveStream with managed models. IAM needs bedrock:AgenticRetrieveStream, bedrock:Retrieve, bedrock:GetDocumentContent and bedrock:InvokeModelWithResponseStream. Guardrails on the agent support BLOCK only, not MASK. Do this before you wire a custom generator.

  1. Pick eu-central-1 (Frankfurt) or eu-west-1 (Ireland). eu-west-2 (London) is not an EU member state.
  2. Create a managed knowledge base. Leave embeddings and rerank on MANAGED unless you have a reason to bring your own model.
  3. Attach a data source. S3 in the same account and Region is the shortest path. SharePoint, Confluence, Google Drive and OneDrive need their connector grants.
  4. Start ingestion. At most 50 jobs run at once. Raw data must stay under 10 TB per knowledge base.
  5. Test Retrieve first. One hybrid search without planning shows whether chunks and metadata are right before the agent spends iterations.
  6. Turn on AgenticRetrieveStream with foundationModelType and rerankingModelType set to MANAGED. Set maxAgentIteration on purpose. The floor is 2.
  7. Attach a guardrail with action BLOCK if you need to filter inputs or answers. AWS applies guardrails to the input and the generated response, not to the retrieved references.
import boto3

client = boto3.client("bedrock-agent-runtime", region_name="eu-central-1")

events = client.agentic_retrieve_stream(
    messages=[
        {
            "role": "user",
            "content": {"text": "Which approval applies to the data processing agreement?"},
        }
    ],
    retrievers=[
        {
            "configuration": {
                "knowledgeBase": {"knowledgeBaseId": "KB12345678"}
            }
        }
    ],
    agenticRetrieveConfiguration={
        "foundationModelType": "MANAGED",
        "rerankingModelType": "MANAGED",
        "maxAgentIteration": 2,
    },
    generateResponse=True,
)

for event in events.get("stream", []):
    if "responseEvent" in event:
        print(event["responseEvent"].get("text", ""), end="")

That is the shape in the API reference. A call may list up to five retrievers, each pointing at one managed knowledge base. retrievalConfigs for AgentCore Memory accepts at most one entry with at most five metadataFilters. sessionBinding and a history you send in messages cannot appear together.

If you want to own the generator, set generateResponse to false and send the hits to Claude on eu.anthropic.claude-sonnet-5 as search_result blocks. Planning and hybrid search stay on AWS. Wording and citations stay on Claude. A gateway and budgets in front, for example LiteLLM in your own EU network, still help once several apps share the same index.

When should you host the index yourself?

Host it yourself when metadata filters, ACLs or the vector store are a hard requirement, or when agentic retrieval is not the loop you want. Stay on the managed knowledge base when connectors, the service-managed embedder and the planning loop are work you do not want to run.

Self-hosting means chunks in OpenSearch, pgvector or Qdrant in Frankfurt, embeddings through Titan V2 in eu-central-1, Claude only as the generator on the EU geo profile. The search results API wants source, title and an array of text blocks. One block is the smallest citable unit. For sentence-level citations, split each chunk into its own block. Images are not allowed inside search_result.

from anthropic import AnthropicBedrock

client = AnthropicBedrock(aws_region="eu-central-1")

knowledge_base_tool = {
    "name": "search_knowledge_base",
    "description": "Search the internal index",
    "input_schema": {
        "type": "object",
        "properties": {"query": {"type": "string"}},
        "required": ["query"],
    },
}

# After the tool call, return hits in this shape:
tool_result = {
    "type": "tool_result",
    "tool_use_id": "toolu_01example",
    "content": [
        {
            "type": "search_result",
            "source": "kb://dpa-2026",
            "title": "Data processing",
            "content": [{"type": "text", "text": "The DPA covers every prompt that includes personal data."}],
            "citations": {"enabled": True},
        }
    ],
}

This path needs evals, tracing and a budget. Langfuse self-hosted in the EU keeps traces on your network. The LLMOps operating model splits retrieval precision from end-to-end correctness. Bedrock model evaluation bills judge tokens at on-demand rates and Knowledge Base usage on top.

Google Cloud is the alternative when your estate already sits on GCP. The RAG Engine overview lists europe-west3 (Frankfurt) and europe-west4 (Eemshaven) as GA. Claude on the eu endpoint remains the residency control for the generator; setup is in Claude on Vertex AI in Europe. The Anthropic API with inference_geo: "us" is not an EU setup, even if the index is in Frankfurt. The provider comparison with Foundry and the first-party API is in Claude GDPR in Europe.

We recommend the managed knowledge base in Frankfurt when SharePoint or Confluence is the source and nobody wants to run a vector cluster. We recommend your own index when you must filter fields, isolate tenants or swap the embedder. Do not mix the two in one request: agentic retrieval accepts managed knowledge bases only.

FAQ

What does agentic RAG cost on Bedrock in the EU?

AWS’s example prices 50 GB plus 100,000 standard searches at 350 USD a month. The same volume as agentic retrieval with two underlying Retrieve calls each lands at 850 USD. Claude tokens for the answer are extra once you pick your own model. Regional endpoints from Sonnet 4.5 cost 10 percent more than global.

Managed Knowledge Base or a self-hosted index?

The managed knowledge base when you want connectors, service-managed embeddings and agentic retrieval. Your own index when ACLs, metadata filters or a specific vector store are mandatory. Agentic retrieval does not run on customer-managed knowledge bases.

AgenticRetrieveStream or RetrieveAndGenerate?

AgenticRetrieveStream plans, retrieves iteratively and cites, only on managed knowledge bases. RetrieveAndGenerate retrieves once and generates, only on customer-managed knowledge bases. AWS blocks RetrieveAndGenerate on the managed index. For a single hybrid search without planning, stay on Retrieve.

Do prompts stay in the EU if Claude answers through the Anthropic API?

No. inference_geo is us or global. Workspace geo is us only. If you keep the index in Frankfurt and send chunks to api.anthropic.com, the prompt leaves the Union. Use eu.anthropic.claude-sonnet-5 on bedrock-runtime or Google’s eu endpoint.

Which quotas apply to a managed knowledge base?

10,000 knowledge bases per account per Region, 200 data sources, 50 concurrent ingestion jobs, 10 TB of raw data, 10,000 characters per query, 600 Retrieve RPM per knowledge base, 300 AgenticRetrieveStream RPM per account. The first value and the last two are adjustable.

Is a knowledge base in London enough for GDPR residency?

London (eu-west-2) appears in the AWS Europe list and is not an EU member state. For residency inside the Union pick Frankfurt or Ireland. Structured-data knowledge bases also exist in Paris and Zurich. Zurich is Switzerland, also outside the Union.

Sources

  1. AWS: Managed Knowledge Bases regions (8 October 2026)
  2. AWS: Managed vs customer-managed Knowledge Bases (8 October 2026)
  3. AWS: Agentic retrieval (8 October 2026)
  4. AWS: Service quotas for managed knowledge bases (8 October 2026)
  5. AWS: Supported models and Regions for Knowledge Bases (8 October 2026)
  6. AWS: RetrieveAndGenerate (8 October 2026)
  7. AWS: Bedrock pricing (8 October 2026)
  8. AWS: Claude Sonnet 5 model card (8 October 2026)
  9. Anthropic: Search results for RAG (8 October 2026)
  10. Anthropic: Pricing (8 October 2026)
  11. Anthropic: Data residency (8 October 2026)
  12. Google Cloud: RAG Engine overview (8 October 2026)

Related guides