Anthropic releases Claude Opus 5.5. Pricing drops 40%, cache costs cut in half.
Anthropic has launched Claude Opus 5.5 with improved performance across benchmarks, 40% lower pricing, and a 60% reduction in prompt cache read costs, making the model more accessible for production deployments.
September 28, 2026

Claude Opus 5.5 lists at $4.00 per million input tokens and $20.00 per million output tokens, according to Anthropic's own announcement. That is the number to sit with before anything else, because Anthropic is not calling this a minor version bump. The company is framing it as a full swap-in replacement for Claude Opus 5, priced roughly 40% lower on the standard rate card and, per Anthropic, cutting cache read costs by 60% on top of that. For a model in Anthropic's top tier, that is a steep price cut to make in one release.
What a cache read cost actually is
If you have never touched the Claude API directly, "cache read cost" sounds like internal plumbing. It is not. Think of it like a reference librarian who has to re-read your entire question every time you ask a follow-up, even if 90% of that question is the same system prompt, the same document, the same instructions you sent five minutes ago. Prompt caching is Anthropic letting the librarian keep a bookmark. Instead of billing you full price to re-process content the model has already seen in the same session, a cache read charges a fraction of the input rate for anything that has not changed.
This matters most for workflows that repeat a large, stable context with small variable inputs on top: a coding agent working through the same codebase across dozens of tool calls, a document analysis pipeline that re-references the same 40-page contract, or a customer support bot that reloads the same product knowledge base for every ticket. In those cases, cache reads can make up the majority of total token volume. A 60% cut to that specific line item is not cosmetic. For a team running Claude Code against a large repository all day, cache reads are often the single biggest cost driver, bigger than the fresh input tokens or the output tokens combined.
The real cost of switching, not just the sticker price
The headline number is the price cut. The number that actually determines whether a team benefits from it is the switching cost, and that number is never zero.
Start with the obvious part: if you are already calling Claude through the API and your integration references a model string rather than a pinned snapshot, the price drop applies the moment Anthropic routes your traffic to Claude Opus 5.5. No code change, no migration. That is the best case and it is real for a meaningful share of API users.
The less obvious part is everything downstream of "the model changed." A prompt tuned against Claude Opus 5's quirks, its exact phrasing tendencies, its specific failure modes on edge cases, is not guaranteed to behave identically on 5.5. Anthropic says performance improved, but "improved" is an aggregate claim across benchmarks, not a guarantee that your particular chain of prompts, your particular JSON schema enforcement, or your particular tone requirements survive the swap unchanged. Teams running Opus in production for anything customer-facing should budget for a regression pass: running the old test set against the new model and diffing outputs, not just trusting the price sheet.
Then there is the human cost. Someone has to notice the release, read the changelog, decide whether the cost savings justify a re-test cycle, and get sign-off from whoever owns the budget line. For a two-person startup that is an afternoon. For a team with a formal model evaluation process and a compliance review, that is a sprint. The 40% price cut is instant. The confidence to rely on it in production is not.
40%
Anthropic's stated price reduction for Claude Opus 5.5 versus its prior Opus tier
A support team's Tuesday migration
Picture a 15-person customer support operation that routes escalated tickets through Claude for draft responses, using a system prompt that loads the company's 60-page policy document as cached context on every call. Before this release, they were running roughly 2 million output tokens a day through Claude Opus 5 at $25 per million, plus cache reads on that 60-page document firing thousands of times a day.
Their engineer swaps the model identifier in the API call from claude-opus-5 to claude-opus-5-5, redeploys the ticket-routing service, and watches the next hour of traffic. Output cost per ticket drops immediately because the per-token output rate fell from $25 to $20 per million. The bigger shift shows up in the cache metrics: the 60-page policy document, previously billed at a reduced cache rate on every repeat call, now bills at 60% less than that already-reduced rate. Across a day with 8,000 tickets, the compounding effect on cache reads outweighs the output savings.
What they do not skip is the verification step. Before rolling the new model to 100% of traffic, they route 10% of tickets through 5.5 for a week, have a human reviewer spot-check 50 of those drafts against the same 50 prompts run through the old model, and only flip the rest once nothing has visibly regressed in tone or accuracy. That is the part of "just switch the model string" that the pricing announcement does not mention.
Why Anthropic is cutting this deep, this fast
Price cuts on flagship models used to be rare events, spaced months apart, framed as a reward for efficiency gains in training or inference. That cadence has collapsed. Anthropic released Claude Opus 5 in July, Claude Opus 5.5 arrives roughly two months later at a meaningfully lower price, and the gap between those two releases tells you more about the competitive environment than about any specific technical breakthrough at Anthropic.
Look at what shipped around the same window: Gemini 3.7 Flash at $0.75/$3.75 per million tokens, Grok 4.7 at $2.00/$6.00, GLM-5.3-Flash undercutting everyone at $0.075/$0.25. None of those are direct Opus-tier competitors on raw capability, but they set the floor that buyers now compare against when a procurement team asks "why are we paying five times more for Claude." Anthropic's answer used to be "because it's better." Cutting Opus pricing by 40% while also compressing cache costs is Anthropic answering a second question buyers are now asking: "why are we paying five times more for Claude, on top of paying again every time it re-reads the same 60 pages."
The cache read discount specifically targets the workflows Anthropic has been courting hardest: coding agents, long-running assistants, and enterprise tools that hold large stable context. Those are exactly the use cases where a competitor's cheaper-but-less-capable model becomes tempting if the cost delta gets wide enough. Anthropic is not just competing on intelligence anymore. It is competing on the shape of its bill.
Switching your integration to Claude Opus 5.5
- Check your current model reference in code. If you call a versioned snapshot (something like
claude-opus-5-20260701) rather than a rolling alias, you will need to update the string manually. - Pull your last 90 days of API logs, if you log them, to estimate your actual split between fresh input tokens, cache reads, and output tokens. This tells you whether the 60% cache discount or the 40% base rate cut matters more for your bill.
- Update the model identifier in a staging environment first, not production. Anthropic's own docs at docs.claude.com list the current model strings and any deprecation timeline for Opus 5.
- Run your existing prompt test set, if you have one, against both models side by side. If you do not have one, this is the moment to build even a small one: 20 representative prompts with expected output shapes.
- Route a small percentage of live traffic to Claude Opus 5.5 for a defined window before full rollout. A day is enough for high-volume workflows; a week is safer for anything customer-facing.
- Confirm your billing dashboard shows the new per-token rates applied to the new model, not the old rates carried over by a cached pricing config on your end.
Verification test: pull one week-old API call, resend the identical prompt and context through Claude Opus 5.5, and compare the invoiced cost for that single call against what the same call cost under Claude Opus 5. If the delta roughly matches the 40% input/output reduction plus a larger drop on any cached context, the migration took.
For teams weighing Opus 5.5 against other frontier options right now, the comparison worth reading is Claude vs Gemini, and if you are deciding between Anthropic's own tiers, Claude's tool page tracks the current lineup. Anthropic's announcement is posted directly at anthropic.com/claude-opus-5-5. For more on how fast this pricing cycle has been moving across the industry, see our earlier look at separating LLM hype from reality.
TL;DR
Claude Opus 5.5 replaces Claude Opus 5 at $4.00/$20.00 per million tokens, roughly 40% cheaper on Anthropic's stated numbers, with a 60% cut to cache read costs that matters most for coding agents and any workflow that repeats large stable context. If you're on the API, swap the model string in staging, re-run your test prompts, and roll out gradually rather than trusting the price cut alone to mean nothing else changed.
Tools mentioned in this article
Some links in this article are affiliate links. Learn more.