Claude Opus 5.5 vs Opus 5: Performance Changes and API Cost Savings
Compare Claude Opus 5.5 and Opus 5 on API pricing, cache costs, effort settings, context, tool use, and the real cost of migration

Claude Opus 5.5 replaces Opus 5 as Anthropic’s current Opus model, but the upgrade is not a simple model-name swap. Published token prices are lower, the default effort level changes, adaptive thinking is always on, and several API patterns must be updated.
For teams deciding whether to migrate, the central finding is this: Opus 5.5 is 20% cheaper at base input and output token rates, but a production bill can change by more or less than 20% depending on cache reads, thinking, effort, retries, and tool use.
Specification comparison
| Feature | Claude Opus 5.5 | Claude Opus 5 | |---|---:|---:| | Input price | $4 / million tokens | $5 / million tokens | | Output price | $20 / million tokens | $25 / million tokens | | Cache read price | $0.20 / million tokens | $0.50 / million tokens | | Context window | 1 million | 1 million | | Maximum standard output | 128,000 | 128,000 | | Thinking | Adaptive, always on | Adaptive | | Default effort | Medium | High | | Knowledge cutoff | June 2026 | May 2026 | | Status | Latest | Legacy, still active |
The context and standard output limits are unchanged. The clearest direct saving is price: both base input and output rates fall by one fifth. Cache reads fall by 60%, which can make a larger difference in workflows that repeatedly reuse a screenplay, story bible, system prompt, or production schema.
Why Anthropic also says “40% less to run”
Anthropic’s launch material says Opus 5.5 costs an estimated 40% less to run than Opus 5 for typical workloads, while the listed input and output prices are only 20% lower. These claims measure different things. The 20% figure compares base token rates directly. The larger estimate can incorporate workload behavior and cheaper cache reads.
Editors should not rewrite the 40% estimate as a universal customer saving. A short, uncached generation with similar token use should be close to the 20% rate reduction. A cache-heavy agent may save more. A high-effort workflow that produces substantially more thinking or output tokens may save less than expected.
Example cost comparison
Consider a screenplay-planning task billed for 75,000 input tokens and 25,000 output tokens, with no cache discount or paid server-side tools:
- Opus 5.5 input: 0.075 × $4 = $0.30
- Opus 5.5 output: 0.025 × $20 = $0.50
- Opus 5.5 total: $0.80
- Opus 5 total: 0.075 × $5 + 0.025 × $25 = $1.00
The direct saving is $0.20, or 20%. This is an arithmetic example, not a promise that every screenplay job uses those token volumes.
Migration changes that affect applications
Developers should review at least five areas:
- Change the model ID to
claude-opus-5-5on the Claude API. - Remove requests that disable thinking or set manual thinking budgets.
- Set
effortexplicitly and rerun quality, latency, and cost evaluations. - Replace forced
tool_choicepatterns that useanyor a named tool. - Read response blocks by type and preserve thinking blocks during tool loops.
Computer-use integrations and interfaces that show progress between tool calls require additional attention. Anthropic’s migration guide documents platform-specific tool changes and explains how progress text may arrive through thinking blocks.
Which model is better for creative production?
For a new implementation, Opus 5.5 is the sensible starting point because it is the current model, has lower published rates, and receives the updated behavior and prompting guidance. Existing Opus 5 applications should migrate only after running their own evaluation set.
A useful creative evaluation set might include:
- preserving character facts across a long script;
- finding timeline contradictions;
- turning scenes into a fixed shot-list schema;
- revising dialogue without changing plot facts;
- maintaining content ratings and production constraints;
- recovering correctly after a failed tool call.
After text planning, Elser AI can turn selected concepts into images, storyboards, voices, sound, and video tests. Teams producing longer projects can try the ElserStudio desktop workflow to organize scripts, reusable assets, shot generation, and composition in one local-first workspace.
Verdict
Opus 5.5 offers a clear base-price reduction and a much larger cache-read reduction, but it also introduces behavior and API changes that require testing. The right migration decision should be based on cost per accepted result—not cost per token alone.
Methodology
Prices and technical limits were checked against Anthropic’s model and migration documentation on October 9, 2026. Cost examples use the published rates and explicitly stated token assumptions. No independent performance benchmark is claimed.








































































































