LLM Observability for Claude: EU Stacks Compared
LLM observability for Claude in the EU: Bedrock invocation logging and CloudWatch, Vertex AI logging to BigQuery, OpenTelemetry and Langfuse compared.
TL;DR
LLM observability for Claude in the EU needs three layers: cloud metrics for latency, tokens and errors, a request log that you control, and traces with evals for quality. On Bedrock, invocation logging is off by default and CloudWatch keeps logs forever unless you set retention. Vertex logs Claude to BigQuery in Preview. Langfuse covers evals.
What should you observe when Claude runs in production?
LLM observability means answering five questions per request: how long did it take, how many tokens did it burn, what did it cost, did it fail, and was the answer any good. The first four come almost free from your cloud. Quality needs traces, scores and human review. Prompts that contain personal data turn every log into a GDPR record.
We split the signals like this:
| Signal | Where it comes from | Contains personal data? |
|---|---|---|
| Latency, time to first token | Cloud metrics (InvocationLatency, TimeToFirstToken on Bedrock) |
No |
| Input, output and cache tokens | Cloud metrics and invocation logs | No |
| Cost per team or app | Token counts times price, grouped by caller | No |
| Errors and throttles | Cloud metrics, client error codes | Rarely |
| Prompt and completion text | Invocation logs, traces | Often |
| Quality scores, evals | Langfuse or a similar trace tool | Depends on what you score |
The split matters for data protection. Metrics are aggregates without content, so they can stay on longer retention and wide access. Prompt logs are content. Treat them like a customer database: short retention, restricted readers, a documented purpose in your records of processing. If you cannot name the purpose, do not switch on text logging.
Claude Code is a separate source with its own telemetry. The Claude Code monitoring docs require CLAUDE_CODE_ENABLE_TELEMETRY=1 and export metrics such as claude_code.token.usage and claude_code.cost.usage every 60 seconds, logs every 5 seconds. Prompt text is not logged unless you set OTEL_LOG_USER_PROMPTS=1. We keep that default. The setup for the Bedrock side is in our Claude Code on Bedrock guide.
How do Bedrock invocation logging and CloudWatch compare to the alternatives?
Bedrock gives you metrics automatically and a full request log on request. Vertex gives you a dashboard plus a Preview log to BigQuery. OpenTelemetry gives you portable traces you send anywhere. Langfuse adds evals on top. For an EU company the decisive columns are where the text lands, how long it stays and who can unmask it.
| Option | Where data sits | EU location | Default retention | PII masking | Cost model | Evals |
|---|---|---|---|---|---|---|
| Bedrock invocation logging + CloudWatch Logs | Your log group or S3 bucket, same account and Region as the call | Any EU Region you call Bedrock from | Indefinite until you set 1 to 3653 days | CloudWatch data protection at ingestion | 0.50 USD/GB ingest, 0.12 USD/GB masking (US East list) | None |
| Vertex request-response logging | Your BigQuery table | EU multi-region or a single EU region |
Table never expires unless you set it | None in the logging feature | BigQuery storage; sampling 0 to 1 | None |
| Cloud Logging (audit and app logs) | Log bucket in your project | Bucket location you choose | 30 days, configurable 1 to 3650 | Redact in the app | 0.50 USD/GiB, 50 GiB free per project | None |
| OpenTelemetry + your backend | Wherever your collector exports | Yours to decide | Yours to decide | Content attributes are opt-in | Your infrastructure | Backend dependent |
| Langfuse (EU cloud or self-hosted) | Langfuse EU cloud in Ireland, or your own cluster | AWS eu-west-1 or your VPC |
30 days Hobby, 90 Core, 3 years Pro | Client-side mask function in the SDK | Free tier, 29 USD, 199 USD; 8 USD per 100k units | LLM-as-a-judge, datasets, annotation queues |
Prices are list prices from the CloudWatch pricing page, which quotes US East (N. Virginia) and says regional prices differ, the Google Cloud Observability pricing page and the Langfuse pricing page. Check your EU Region before budgeting.
Our position: start with the native layer of the cloud you already use for Claude. It is inside your existing DPA with AWS or Google, it needs no new subprocessor, and it already has the token counts. Add OpenTelemetry for application traces. Add Langfuse only when somebody will actually run evals. Teams that start with a trace tool and no cloud metrics miss throttling, which only the cloud sees.
How do you enable Bedrock invocation logging with the AWS CLI?
Model invocation logging is disabled by default and is set per Region. Once enabled, it captures the full request, response and metadata for Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream on the bedrock-runtime endpoint. Destinations must be in the same account and Region. Set log group retention first, because CloudWatch keeps logs forever by default.
The invocation logging guide lists what you need before the API call: a log group, an IAM role that bedrock.amazonaws.com can assume with logs:CreateLogStream and logs:PutLogEvents, and optionally an S3 bucket for bodies above 100 KB.
- Create the log group in the Region you call Bedrock from, for example Frankfurt.
- Set retention. The put-retention-policy reference accepts fixed values from 1 to 3653 days. We use 30 for prompt text.
- Create the IAM role with the trust policy from the AWS guide, scoped to your account and Region.
- Enable logging with text delivery on and the modalities you do not need off.
- Read the configuration back and send one test request.
aws logs create-log-group --log-group-name /bedrock/invocations --region eu-central-1
aws logs put-retention-policy --log-group-name /bedrock/invocations --retention-in-days 30 --region eu-central-1
aws bedrock put-model-invocation-logging-configuration --region eu-central-1 --logging-config '{"cloudWatchConfig":{"logGroupName":"/bedrock/invocations","roleArn":"arn:aws:iam::123456789012:role/BedrockInvocationLogging","largeDataDeliveryS3Config":{"bucketName":"example-bedrock-logs-eu","keyPrefix":"large"}},"textDataDeliveryEnabled":true,"imageDataDeliveryEnabled":false,"embeddingDataDeliveryEnabled":false,"videoDataDeliveryEnabled":false}'
aws bedrock get-model-invocation-logging-configuration --region eu-central-1
Each log entry carries identity.arn, modelId, inputTokenCount and outputTokenCount, and an optional requestMetadata object you set per call. That gives you cost per team without a separate tool: group by identity.arn in Logs Insights. Invocation logging also emits its own delivery metrics, such as ModelInvocationLogsCloudWatchDeliveryFailure. Alarm on that one, or you will discover a broken IAM role during an audit.
Two traps. First, calls to bedrock-mantle are not captured. If you use single-Region Sonnet 5 through mantle, you have metrics and CloudTrail but no invocation log. Second, the configuration is per Region. With the EU geo profile you call from one source Region; configure every source Region your apps use. The model ID choices behind that are in our Claude on AWS Bedrock EU guide.
For personal data, attach a CloudWatch Logs data protection policy before you enable text logging. It masks matches at ingestion. Events written before the policy existed stay unmasked. Only principals with logs:Unmask see raw values, which gives you an access control you can document.
What does Google Cloud offer for Claude on Vertex AI?
Google gives you a prebuilt model observability dashboard for managed models, partner models included, plus request-response logging to BigQuery. The logging is Preview, supports Claude through rawPredict and streamRawPredict, and for Anthropic models can only be configured through the REST API. You choose a sampling rate between 0 and 1.
The request-response logging page uses setPublisherModelConfig with publisher set to anthropic. The body sets enabled, samplingRate, a BigQuery outputUri and optionally enableOtelLogging, which adds an otel_log column in OpenTelemetry format. Request-response pairs larger than the 10 MB BigQuery row limit are not recorded. A long agent transcript can silently fall out of your log.
Location is your job. Create the dataset in the EU multi-region or a single EU region before you point the config at it. BigQuery documents that EU multi-region data is stored in europe-west1 (Belgium) or europe-west4 (Netherlands). Tables never expire unless you set an expiration, so set one on the dataset. The EU endpoint choices for Claude itself are in our Claude on Vertex AI Europe guide.
One setting deserves a policy. For Claude Mythos Preview, Claude Mythos 5 and Claude Fable 5 on Google Cloud, Vertex can share logged requests and responses with Anthropic in real time once dataSharingEnabledProvider is set to ANTHROPIC under the Advanced AI Safety Addendum. That is a transfer to a second party. If you do not want it, Google documents an organization policy custom constraint that denies the field, and VPC Service Controls block sharing by default.
Cloud Logging is the place for audit logs and your own application logs, not for model payloads. Log buckets default to 30 days and accept 1 to 3650. Ingestion costs 0.50 USD per GiB with 30 days included and 50 GiB free per project per month.
Where do OpenTelemetry and Langfuse fit?
OpenTelemetry is the vendor-neutral way to trace your own application: one span per Claude call, linked to the HTTP request and the tool calls around it. The GenAI semantic conventions define names for model, provider and token usage. Langfuse consumes those traces and adds what the clouds lack: datasets, LLM-as-a-judge scoring and human annotation.
The GenAI semantic conventions moved to their own repository and still carry status Development. Expect attribute names to change. For Claude, gen_ai.provider.name is anthropic, or aws.bedrock when you call through Bedrock. Token counts go into gen_ai.usage.input_tokens and gen_ai.usage.output_tokens, with separate cache read and write attributes. The important part for GDPR: gen_ai.input.messages and gen_ai.output.messages are Opt-In. A conforming instrumentation does not record prompt text unless you ask for it.
We keep Langfuse short here and stick to the EU facts. The EU cloud runs at cloud.langfuse.com in Ireland on AWS eu-west-1. Self-hosting is free. The OTLP endpoint is /api/public/otel over HTTP/JSON or HTTP/protobuf. gRPC is not supported, so an exporter configured with OTEL_EXPORTER_OTLP_PROTOCOL=grpc will not reach it. Masking runs client-side in the SDK before export.
A reference stack we would build for an EU company on Bedrock:
- CloudWatch alarms on
InvocationThrottles,InvocationServerErrorsandTimeToFirstTokenperModelId. - Invocation logging in each EU source Region, 30-day retention, data protection policy, large bodies to an EU S3 bucket with a lifecycle rule.
- OpenTelemetry in the application with content attributes off, exporting to a collector in your EU VPC.
- Langfuse self-hosted, or the EU cloud under a DPA, fed by a sampled slice of traces for evals.
- Claude Code telemetry to the same collector, metrics only.
On Google Cloud, swap step 1 for the model observability dashboard and step 2 for BigQuery logging with sampling.
Run it yourself or outsource?
Running LLM observability in-house is mostly recurring work, not setup. Somebody owns alert thresholds as traffic shifts, retention and deletion against the records of processing, dashboard upkeep, and the eval loop. The setup above takes a competent platform engineer days. The ongoing part is what teams underestimate.
What it costs in-house, by role (our estimate, not a sourced benchmark):
| Task | Who | Rhythm |
|---|---|---|
| Alert tuning (throttles, latency, error rate) | Platform or SRE | Weekly at first, then monthly |
| Retention, deletion, access reviews for prompt logs | Platform with the DPO | Quarterly |
| Dashboards per team and cost allocation | Platform or FinOps | Monthly |
| Eval datasets, judge prompts, annotation | Product owner and domain experts | Continuous |
| Upgrades of collector and Langfuse when self-hosted | Platform | Per release |
Outsourcing makes sense when you have several Claude apps but no SRE function, when on-call for a log pipeline is out of the question, or when you need audit-ready retention evidence fast. It makes little sense for evals: a provider cannot judge whether an answer is right for your domain. Keep that in-house.
What to demand from a managed service provider for this topic:
- An SLA on log delivery and alert response, not only on dashboard uptime.
- A DPA that names every place prompt text is stored, plus the subprocessor list, including any observability SaaS they add.
- Access model: least privilege, unmasking only through a named role, every access logged.
- Retention that matches your records of processing, with deletion you can verify.
- Exit: logs and dashboards stay in your own AWS or Google Cloud account, exportable, with no proprietary agent you cannot remove.
FAQ
How much does LLM observability for Claude cost on AWS?
Metrics in the AWS/Bedrock namespace come with the service. Invocation logs cost CloudWatch Logs ingestion, listed at 0.50 USD per GB in US East, plus 0.03 USD per GB archived and 0.12 USD per GB if a data protection policy scans them. 5 GB are free. EU Region prices differ.
CloudWatch vs Langfuse: which one do I need?
Both, for different jobs. CloudWatch sees what only AWS sees: throttles, server errors, time to first token and a log tied to the IAM identity. Langfuse adds evals, datasets and human review. If nobody will score answers, CloudWatch alone is enough.
Does Bedrock log my prompts by default?
No. Model invocation logging is disabled by default. Once you enable it with text delivery, the full request and response bodies up to 100 KB go into your CloudWatch log group, and larger bodies go to S3. Retention is indefinite until you set it.
Does invocation logging capture bedrock-mantle calls?
No. AWS states that invocation logging covers only the bedrock-runtime endpoint. Calls through bedrock-mantle, including the Anthropic Messages API there, are not captured. You still get CloudWatch metrics and CloudTrail for mantle.
Can Google share my Claude logs with Anthropic?
Only if someone enables it. For Claude Mythos Preview, Mythos 5 and Fable 5 on Google Cloud, setting dataSharingEnabledProvider to ANTHROPIC shares logged requests and responses in real time. An organization policy custom constraint or a VPC Service Controls perimeter prevents it.
Do OpenTelemetry GenAI conventions record prompt text?
Not by default. In the GenAI semantic conventions, gen_ai.input.messages and gen_ai.output.messages are Opt-In attributes. The conventions are still in Development status, so pin your instrumentation version and expect renames.
Sources
- AWS: Monitor model invocation using CloudWatch Logs and Amazon S3 (1 October 2026)
- AWS: Monitor bedrock-runtime inference using CloudWatch metrics (1 October 2026)
- AWS CLI: put-model-invocation-logging-configuration (1 October 2026)
- AWS: CloudWatch Logs data protection (masking) (1 October 2026)
- AWS: Working with log groups (retention) (1 October 2026)
- AWS CLI: logs put-retention-policy (1 October 2026)
- Amazon CloudWatch pricing (1 October 2026)
- Google Cloud: Log and share requests and responses (1 October 2026)
- Google Cloud: Monitor models (1 October 2026)
- Google Cloud: BigQuery locations (1 October 2026)
- Google Cloud: Configure log buckets (1 October 2026)
- Google Cloud Observability pricing (1 October 2026)
- OpenTelemetry GenAI semantic conventions (1 October 2026)
- Claude Code: Monitoring usage (1 October 2026)
- Langfuse: Data regions (1 October 2026)
- Langfuse pricing (1 October 2026)
- Langfuse: OpenTelemetry integration (1 October 2026)
- AWS: Claude Sonnet 5 model card (1 October 2026)