LLMOps for Claude: Operating Model, Roles and SLOs
LLMOps for Claude in production: how it differs from MLOps, the lifecycle from prompt versioning to model retirement, recommended SLOs, tooling and roles.
TL;DR
LLMOps is the operating discipline for applications built on a hosted model such as Claude. You do not train weights. You version prompts, gate every change on evals, watch latency, errors, throttles and token cost, and migrate before a model retires. Anthropic gives at least 60 days’ notice. Bedrock usually gives 6 months. Plan for both.
What is LLMOps, and how is it different from MLOps?
LLMOps is the set of practices that keeps an LLM application correct, fast, affordable and current once real users depend on it. With Claude via Bedrock, Google Cloud or the Claude API, the model is a rented dependency. Your levers are prompts, context, parameters, routing and the model version you pin. MLOps assumes you own the model.
That shift changes what the team does all week. Nobody retrains Claude. Instead, a vendor ships new versions, retires old ones on a fixed calendar, and changes which request parameters a model accepts. Cost is per token, not per GPU hour, so a longer prompt is a budget decision. Quality is judged on open text, so unit tests are not enough.
| MLOps (own model) | LLMOps (hosted Claude) | |
|---|---|---|
| What you change | Training data, features, weights | Prompts, context, tools, model ID, parameters |
| Release unit | Model artifact | Prompt version plus pinned model ID |
| Quality check | Holdout metrics (accuracy, AUC) | Eval set graded by code, humans or a second LLM |
| Main cost driver | Training and serving compute | Input, output and cache tokens per request |
| Capacity limit | Your cluster | Provider quotas (RPM, input and output tokens per minute) |
| Lifecycle risk | Data drift | Provider retires the model on a date you do not set |
| Data protection focus | Training data | Prompts and responses in logs |
We treat LLMOps as a narrower job than MLOps, not a bigger one. If you already run SRE and change management, most of it fits into existing processes. The parts that are new are evals, token economics and model retirement.
What does the LLMOps lifecycle for Claude look like?
The lifecycle has six recurring steps: version the prompt, evaluate it, deploy it behind a pinned model ID, monitor it, control cost, and upgrade the model before retirement. Each step needs an owner and a trigger. The step teams skip most often is the last one, and it is the only one with a hard external deadline.
- Version prompts like code. System prompt, tool definitions and few-shot examples live in Git, next to the pinned model ID. A prompt change is a pull request.
- Gate on evals. Anthropic’s eval guidance asks for specific, measurable success criteria and automated grading, and says more questions with slightly noisier automated grading beat fewer hand-graded ones. It also recommends using a different model to grade than the one that produced the output.
- Deploy with a pinned ID. Use a full model or inference profile ID, such as
eu.anthropic.claude-sonnet-5on Bedrock, never an alias you do not control. - Monitor errors, throttles, latency and tokens per model (next section).
- Control cost with prompt caching and budgets. On Claude Opus 5.5, base input is 4 USD per million tokens and a cache hit is 0.20 USD, per the prompt caching page. Standard models read cache at 0.1x the input price.
- Upgrade before retirement. Rerun the full eval set on the replacement, then switch the pinned ID.
Step 6 is where production breaks. The Anthropic deprecations page promises at least 60 days’ notice for publicly released models. On 30 September 2026 Anthropic deprecated claude-sonnet-4-5-20250929, retiring on 30 November 2026, with claude-sonnet-5-5 as replacement. Those dates apply to the Claude API. Bedrock and Google Cloud run their own schedules. The Bedrock lifecycle policy uses Legacy periods of 6 months or 45 days, and a Legacy model can lose access after 15 days of inactivity.
An upgrade is not a string swap. On Claude 4.7 and later, setting temperature, top_p or top_k to a non-default value returns HTTP 400. A prompt that worked for a year can fail on the first request to the new model. This is exactly what the eval run before the switch should catch.
Which SLOs should an LLM app on Claude have?
Set four SLOs per application: availability, latency, cost per request and eval pass rate. The targets below are our recommendations as starting points, not provider guarantees. Calibrate them on two weeks of your own baseline. Providers publish limits and metrics, not SLOs for your app, so the error budget is yours to define.
| SLO (our recommendation) | Starting target | Measure on Bedrock | Measure on Google Cloud |
|---|---|---|---|
| Availability: requests that succeed after retries | 99.5% per 30 days | InvocationServerErrors, InvocationThrottles vs Invocations |
prediction/online/error_count vs response_count |
| Time to first token, streaming chat, p95 | under 3 s | TimeToFirstToken |
First token latency on the model observability dashboard |
| Full response latency, p95 | set per use case | InvocationLatency |
prediction/online/prediction_latencies |
| Cost per request, p50 | baseline plus 20% | Invocation log token counts | Request-response log, token_count metrics |
| Eval pass rate on the golden set | no drop vs current version | Your eval pipeline | Your eval pipeline |
The Bedrock runtime metrics under AWS/Bedrock include TimeToFirstToken for streaming calls and EstimatedTPMQuotaUsage. AWS warns that the latter is an approximation and should not be the only input for capacity planning. Throttles also do not count as invocations or errors, so compute availability with throttles as a separate term. On Google Cloud, the model observability dashboard shows QPS, token throughput and first token latency for managed partner models including Claude.
Throttling deserves its own alert. On Bedrock, quota is reserved at the start of each request as input tokens plus max_tokens. Output tokens then burn quota at 10x for Claude Opus 5.5, Sonnet 5, Opus 5 and Fable 5.1, per the token counting page. A generous max_tokens default can throttle a service that is nowhere near its real usage.
aws cloudwatch put-metric-alarm --region eu-central-1 --alarm-name claude-throttles --namespace AWS/Bedrock --metric-name InvocationThrottles --dimensions Name=ModelId,Value=<model-id-as-shown-in-cloudwatch> --statistic Sum --period 300 --evaluation-periods 1 --threshold 10 --comparison-operator GreaterThanThreshold --alarm-actions <sns-topic-arn>
On the Claude API, the 429 for a rate limit carries a retry-after header. The 429 for a reached monthly spend cap does not, and retries fail until access resumes, per the rate limits page. Alert on error.details.error_code = enforced_spend_limit_reached separately. It is a budget incident, not a capacity incident.
What does the LLMOps tooling map look like?
Four tool categories cover most of LLMOps for Claude: a gateway for keys, routing and budgets; tracing for per-request visibility; an eval runner; and cost reporting. Start with what the cloud already provides, add OpenTelemetry instrumentation, and add a dedicated tool only when a gap hurts. We keep deep dives in separate articles.
| Category | Built into the platform | Open standard or self-hosted option |
|---|---|---|
| Gateway | Bedrock IAM and inference profiles; Google Cloud quotas | LiteLLM in your own EU network |
| Request logs | Bedrock invocation logging (CloudWatch Logs, S3); Google Cloud request-response logging (BigQuery) | OTLP to your collector |
| Tracing | CloudWatch, Cloud Monitoring | OpenTelemetry GenAI conventions; Langfuse via OTLP |
| Evals | Bedrock model evaluation | Your own harness in CI |
| Cost | Cost Explorer, Cloud Billing; Anthropic Usage and Cost API | Gateway spend logs |
Request logs are where GDPR meets LLMOps. Bedrock invocation logging is off by default. When enabled, it captures full request and response bodies, inline up to 100 KB and in S3 above that, and destinations must be in the same account and Region. Google Cloud request-response logging supports Claude in Preview, writes to BigQuery and takes a sampling rate between 0 and 1. Both logs will contain whatever personal data users type. Set retention and access before you switch them on.
For portable tracing, the OpenTelemetry GenAI conventions are still at status Development, but they already define an Anthropic profile. One detail matters for cost dashboards: Anthropic’s input_tokens excludes cached tokens, so the convention computes total input as input_tokens + cache_read + cache_write.
from opentelemetry import trace
tracer = trace.get_tracer("claude-app")
def record_usage(span, model: str, usage) -> None:
span.set_attribute("gen_ai.provider.name", "anthropic")
span.set_attribute("gen_ai.request.model", model)
cache_read = usage.cache_read_input_tokens or 0
cache_write = usage.cache_creation_input_tokens or 0
span.set_attribute("gen_ai.usage.cache_read.input_tokens", cache_read)
span.set_attribute("gen_ai.usage.cache_write.input_tokens", cache_write)
span.set_attribute("gen_ai.usage.input_tokens", usage.input_tokens + cache_read + cache_write)
span.set_attribute("gen_ai.usage.output_tokens", usage.output_tokens)
Langfuse accepts OTLP on /api/public/otel over HTTP (JSON or protobuf, not gRPC), so the same spans can feed a self-hosted trace UI. Token prices behind these dashboards are in our Claude API pricing comparison.
Who owns what in an LLMOps team?
Four roles carry LLMOps: a product owner for quality criteria, an application engineer for prompts and evals, a platform engineer for gateway, quotas and logging, and an on-call rotation for incidents. In small teams one person holds two roles. What fails is leaving model retirement and eval maintenance without a named owner.
| Role | Owns | Recurring task |
|---|---|---|
| Product owner | Success criteria, eval acceptance | Approves eval set changes and quality trade-offs |
| Application engineer | Prompts, tools, eval cases | Adds failing production cases to the eval set |
| Platform engineer | Gateway, quotas, invocation logging, IAM | Quota requests, max_tokens tuning, log retention |
| On-call | Availability and latency SLOs | Throttle and 5xx alerts, provider status, rollback |
| Data protection officer | Log content, DPA, subprocessors | Reviews logging scope and retention |
Model retirement belongs to the platform engineer, with the application engineer running the evals. The trigger is the deprecation email or the Legacy flag on a Bedrock model card. For Sonnet 5 that card currently says EOL no sooner than 30 June 2027 with a Legacy period of at least 6 months. Region choices for the same models are in our Bedrock EU guide.
Run it yourself or outsource?
Run LLMOps yourself when Claude powers a core product and you already have on-call and a platform team. Outsource parts when you have one or two Claude applications, no 24/7 rotation, and nobody whose job includes reading deprecation notices. Evals stay with you in both cases, because only your team knows what a correct answer is.
What it costs in-house, as our estimate rather than a measured figure: for one production application, roughly half a platform engineer for gateway, quotas, logging and upgrades, a few hours per week of application engineering for eval maintenance, and a share of an existing on-call rotation. The recurring tasks are fixed: weekly review of throttles and cost per request, monthly eval set refresh, and one model migration project per retirement notice.
Outsourcing to a managed service provider makes sense for the platform layer: on-call for the gateway and quotas, invocation logging setup, and deprecation tracking across Bedrock and Google Cloud schedules. It makes less sense for eval design and prompt changes. A provider that writes your prompts without your eval gate ships changes you cannot judge.
What to demand from a provider for LLMOps specifically:
- SLA tied to your SLOs: response time for throttle and 5xx alerts, with the SLO definitions in the contract, not only “best effort”.
- DPA (Art. 28 GDPR) covering logs: where invocation logs and traces live, retention, and who can read prompt content.
- Subprocessor list, including any tracing or eval SaaS the provider adds.
- Access model: their access runs through roles in your AWS or Google Cloud account, logged in CloudTrail or Cloud Audit Logs, never through shared root keys.
- Deprecation duty: a written commitment to flag every retirement notice and run your eval set on the replacement before the switch.
- Exit: prompts, eval sets, dashboards and alarm definitions in your repository, so a handover does not start from zero.
FAQ
What is the difference between LLMOps and MLOps?
MLOps manages models you train: data pipelines, training runs, model artifacts. LLMOps manages applications on a hosted model you do not train. The work moves to prompt versioning, evals, token cost, provider quotas and migrating before the provider retires a model version.
How much notice do you get before a Claude model is retired?
On the Claude API, Anthropic gives at least 60 days’ notice for publicly released models. Bedrock sets its own dates, with a Legacy period of 6 months or 45 days; most models get 6 months. Google Cloud also sets its own schedule. Check the platform you actually call.
How much does prompt caching save on Claude?
A cache read costs 0.1x the base input price on standard models and 0.05x on Claude Opus 5.5, where input is 4 USD and a cache hit 0.20 USD per million tokens. A 5-minute cache write costs 1.25x. On the Claude API, cache reads also do not count toward input tokens per minute for most models.
Which metrics should we alert on for Claude on Bedrock?
Alert on InvocationThrottles, InvocationServerErrors and TimeToFirstToken, per model ID, in the AWS/Bedrock namespace. Track InputTokenCount, OutputTokenCount and the cache token counts for cost. Treat EstimatedTPMQuotaUsage as an approximation, as AWS itself does.
Are LLM SLOs provided by Anthropic or AWS?
No. Providers publish rate limits, quotas, error codes and metrics. The SLOs for your application, such as 99.5% availability or a time to first token target, are your own definitions. The targets in this article are our recommendations as a starting point.
Should invocation logging be on in production?
Yes, with limits. Without request logs you cannot debug or build eval cases from real failures. But logs contain personal data. Restrict access, set a retention period, document it in your records of processing, and sample on Google Cloud if full logging is not needed.
Sources
- Anthropic: Model deprecations (1 October 2026)
- Anthropic: Rate limits (1 October 2026)
- Anthropic: Prompt caching (1 October 2026)
- Anthropic: Define success criteria and build evaluations (1 October 2026)
- Anthropic: API errors (1 October 2026)
- Anthropic: Usage and Cost API (1 October 2026)
- AWS: Bedrock runtime CloudWatch metrics (1 October 2026)
- AWS: Bedrock model invocation logging (1 October 2026)
- AWS: How tokens are counted in Amazon Bedrock (1 October 2026)
- AWS: Bedrock model lifecycle (1 October 2026)
- AWS: Claude Sonnet 5 model card (1 October 2026)
- Google Cloud: Quotas for Anthropic Claude models (1 October 2026)
- Google Cloud: Log requests and responses (1 October 2026)
- Google Cloud: Monitor models (1 October 2026)
- OpenTelemetry: GenAI semantic conventions (1 October 2026)
- Langfuse: OpenTelemetry integration (1 October 2026)