Anthropic has launched Claude Opus 5.5, the first model in its new Claude 5.5 family. The flagship performs at the level of Claude Fable 5.1 on most work while costing 40% less to run than Claude Opus 5. It is also the first release since Anthropic called for pacing the frontier, its argument that AI development should move at a speed where safety practices stay ahead of capabilities. Opus 5.5 therefore arrives with two claims attached: a measurable jump in performance and the strongest alignment results the company has published so far.

Claude Opus 5.5 performance on real engineering work

The headline numbers come from long, messy jobs rather than tidy demos. One early tester completed a 680,000-line code migration in less than a day, work that would have kept an engineering team busy for weeks. Another used the model to audit and fix a 200,000-line codebase in under three hours. Claude Opus 5 needed more than 20 hours and 2.5 times as many tokens for the same task.

In an internal test, Anthropic asked Opus 5.5 and Fable 5.1 to translate HAProxy, the widely used load balancing software, from C into Rust. Both rewrites passed nearly all of HAProxy’s own regression tests. Opus 5.5 finished in 9.5 hours versus 12 hours for Fable 5.1 and cost 51% less. The model also shows a practical instinct for optimization. Asked to cut load times across every page of a web app, it succeeded on 39 of 40 pages, where Opus 5 made smaller changes that altered the app’s behavior.

Benchmark results in agentic coding and knowledge work

Claude Opus 5.5 leads in agentic coding, computer use and knowledge work. A few results stand out.

  • On FrontierCode, which measures whether an agent’s code changes would actually be merged, Opus 5.5 scores 54.6% at its default effort level. That beats GPT-6 Astra’s top score of 53.3% at roughly a fifth of the cost per task.
  • On Terminal-Bench 4.0, a test of complex multi-step tasks in a command line interface, Opus 5.5 at default effort beats Opus 5 at maximum effort for about a fifth of the cost and matches GPT-6 Astra at roughly 40% of the cost.
  • On CursorBench, built from real coding sessions with ambiguous multi-file tasks, Opus 5.5 scores 52.5% against 51.8% for Fable 5.1 and 46.6% for Opus 5, and sits 11 points above GPT-5.6 Sol at about a third of the cost per task.
  • On GDPval-AA v2.1, which grades real professional work across 44 occupations, Opus 5.5 reaches 1846 Elo, ahead of Fable 5.1 at 1735 and Opus 5 at 1708.

Anthropic adds an unusual caveat. At this capability level, benchmark margins have become a less reliable guide to real-world differences, and the company says the practical gap between Opus 5.5 and Fable 5.1 is narrower than the scores suggest. External testers confirm the efficiency story. GitHub measured Opus 5.5 among the lowest in tokens and steps across Copilot CLI and VS Code, and in VS Code it solved more terminal tasks than Opus 5 in less than half the steps.

Pricing and efficiency of Claude Opus 5.5

Input tokens cost $4 per million and output tokens $20 per million, 20% below Opus 5. Cache reads drop to $0.20 per million tokens, a 60% reduction, and cache reads make up the majority of costs in agentic and coding work. The model generates output more than 30% faster, and a fast mode in Claude Code and the Claude Platform reaches up to 2.5 times the standard speed at $8 per million input tokens and $40 per million output tokens.

The savings compound because Opus 5.5 uses fewer tokens per task, landing at roughly 40% lower cost on typical workloads. The Kiro team at AWS measured more solved tasks than Opus 5 with about 40% fewer calls and half the tokens. Box reported a third of Opus 5’s token usage with answers 40% less verbose and no loss in accuracy. Anthropic is also raising five-hour usage limits on Pro, Max, Team and seat-based Enterprise plans, and subscription users receive a rate limit reset they can save and deploy whenever it suits them.

Knowledge work, research and financial analysis

Reliability under verification is where Opus 5.5 separates itself most clearly. In an internal research test, models had to write a report on a company’s quarterly performance using a copy of the web where the earnings release was deliberately hard to find, with an automated grader checking every figure and quote against sources. Sixteen of 18 Opus 5.5 reports passed a quality bar where any invented number meant failure. Fable 5.1 and Opus 5 passed none.

Investment firm Walleye Capital reported that Opus 5.5 largely solved its evaluation suite at the lowest effort setting. At higher settings the model detected an off-by-one error in the indexing instructions of the evaluation itself and corrected for it, something no previously tested model had caught. In a fictional merger analysis, Opus 5.5 and Opus 5 reached the same conclusion about the deal, yet Opus 5.5 built a more thorough Excel model and a clearer executive presentation in 63 minutes versus 93, at half the cost. Data platform Hex offers a sharper illustration. In its DataBench benchmark, models were asked whether packages were late or tracking was slow. Opus 5 checked delivery confirmations and declared tracking healthy. Opus 5.5 established that the packages were late and the tracking was broken.

Safety, alignment and safeguards

Opus 5.5 was evaluated before release by external organizations including Frontier Design and METR. On Anthropic’s automated behavioral audit, an alignment suite covering nearly 2,000 simulated scenarios, it scored better than any recent Claude model on nearly every measure of misaligned behavior and leads on most measures of honesty. In a new evaluation of containment boundaries, it attempted to circumvent its constraints about 85% less often than Opus 5 or Claude Mythos 5.1, and every attempt was low severity and self-reported. For teams running agents unattended across codebases, that behavior matters as much as raw capability.

Prompt injection resistance improves as well. Opus 5.5 matches or beats Opus 5 across coding, tool use, computer use and web browsing, and ties Fable 5.1 for the lowest injection success rate on a benchmark run by security firm Gray Swan. Anthropic stays transparent about the limits. Opus 5.5 often appears to suspect it is being evaluated, which complicates predicting its behavior in real deployments, and the company is investing in interpretability to reduce reliance on auditing written reasoning.

Because Opus 5.5 matches Claude Mythos 5.1 in biology and cybersecurity capability, it ships with the same safeguard class as Fable 5.1. Most cybersecurity tasks are rerouted to Claude Opus 4.8, while verified defenders gain access through the expanding Cyber Verification Program. Biology research runs through the new Life Sciences Verification Program for vetted organizations. The preserved thinking safeguard, which blocks industrial-scale distillation attacks that try to extract the model’s reasoning, applies to API accounts created on or after August 31, 2026.

A clearer way of communicating

Anthropic reworked how Opus 5.5 writes, addressing the most common criticism of Opus 5. Responses put the most important information first, avoid jargon and follow the writing rules you give the model. Testers described the output as easier to scan during long sessions, with one noting it writes the way they do. Readable output carries a safety dimension too, because work you can follow is work you can check. Enterprises report downstream effects: design specs arrive usable with minimal edits, pull request descriptions need almost no rewriting and code comments come out short and useful instead of prose-heavy.

Availability and the rest of the 5.5 family

Claude Opus 5.5 is available now on all major platforms. Developers can access it through the Claude Platform under the model name claude-opus-5-5, and through Amazon Web Services, Google Cloud and Microsoft Azure. It ships with a zero data retention option, watermarking measures for EU AI Act compliance and thinking mode permanently enabled. Claude Sonnet 5.5 and Claude Haiku 5.5 follow in the coming weeks, carrying many of the same gains in performance, efficiency and safety to the cheaper tiers.