Anthropic introduced Claude Opus 5.5 on September 22, 2026, the first model in its new Claude 5.5 family. The company says Opus 5.5 performs at the level of Claude Fable 5.1 on most tasks while costing 40% less to run than Opus 5 on typical workloads.

It is Anthropic's first release since the company called for pacing the frontier. Opus 5.5 was tested before release by external evaluators including Frontier Design and METR, and posts the strongest scores to date on Anthropic's automated behavioral audit, an alignment test spanning thousands of simulated scenarios.

Confirmed

  • Performance: Anthropic reports Opus 5.5 leading on agentic coding, computer use, and knowledge work benchmarks — 66.4% on Terminal-Bench 4.0, 54.4% on FrontierCode v1.1 (Main), 57.8% on CursorBench 4.0, and 1846 on GDPval-AA v2.1, ahead of both Fable 5.1 and Opus 5 on each.
Opus 5.5 Fable 5.1 Opus 5 GPT-6 Astra GPT-5.6 Sol
Agentic codingTerminal-Bench 4.0¹ 66.4% 55.8% 52.3% 57.9% 37.3%
Agentic codingFrontierCode v1.1 (Main) 54.4% 50.3% 48.0% 53.3% 47.5%
Agentic codingCursorBench 4.0 57.8% 51.8% 46.6% 41.7%
Knowledge workGDPval-AA v2.1 1846 1735 1708 1542 1588
Business workflowsAutomationBench² 40.0% 31.4% 26.9% 41.4% 28.8%
Multidisciplinary reasoningHumanity's Last Exam 67.7%with tools 65.6%with tools 63.6%with tools 57.2%with tools
Agentic scientific researchTerminal-Bench-Science 0.1³ 58.7% 52.6% 29.0% 64.6% 22.4%
Computer useOSWorld 2.0 81.8%partial 80.7%partial 74.0%partial
Visual chart recognitionChartography 89.0%with tools 88.4%with tools 83.4%with tools
Claude Opus 5.5 benchmark comparison — Brocker chart (vendor-reported figures)
Brocker chart reconstruction from Anthropic-reported figures — not a vendor asset. GDPval-AA v2.1 omitted (different scale).
  • Cost and speed: $4 per million input tokens and $20 per million output tokens (20% less than Opus 5); cache reads at $0.20 per million (60% less) and cache writes at $5 per million. Fast mode runs up to 2.5x speed in Claude Code and the Claude Platform at $8/$40 per million. Output generation is more than 30% faster than Opus 5, five-hour usage limits rise on Pro, Max, Team, and seat-based Enterprise plans, and subscribers get a rate-limit reset they can save and use at will.
Prices per 1M tokens Claude Opus 5.5 Claude Opus 5
Cache reads $0.20 $0.50
Input tokens $4 $5
Output tokens $20 $25
Cache writes $5 $6.25
  • Safety: Best scores to date on Anthropic's automated behavioral audit, with the model less likely than recent releases to take hard-to-reverse actions or act outside given boundaries, and more resistant to prompt injection than Opus 5. Alignment testing was broadened to longer tasks, impossible tasks, and scenarios modeled on real incidents. Deployed with safeguards similar to Claude Fable 5.1 given comparable biology and cybersecurity capability; Life Sciences Verification Program applications open, with Cyber Verification Program access expanding in coming weeks.
  • Communication: Writes more naturally, leads with the important information, and produces clearer, easier-to-follow output than Opus 5. Claude Sonnet 5.5 and Claude Haiku 5.5 follow in the coming weeks with many of the same improvements.
  • Availability: Opus 5.5 is now available on all platforms — Amazon Web Services, Google Cloud, and Microsoft Azure — and on the Claude Platform as claude-opus-5-5.

Unknown

  • Independent replication of Anthropic's benchmark scores (Terminal-Bench, FrontierCode, CursorBench, GDPval-AA, AutomationBench, Humanity's Last Exam, Terminal-Bench-Science, OSWorld, Chartography) has not been published.
  • Exact rollout timeline and feature parity for Sonnet 5.5 and Haiku 5.5.
  • Full details of the Opus 5.5 System Card (referenced but not fully reproduced in the announcement).
  • Whether the 40% cost reduction on typical workloads holds across diverse real-world agentic coding patterns beyond Anthropic's test scenarios.
  • Specific safeguards deployed for biology and cybersecurity use cases beyond the Life Sciences and Cyber Verification Programs.

Our take

Opus 5.5 is positioned as a workhorse that closes much of the distance to Anthropic's own Fable 5.1 tier while cutting the bill. The cost drop — lower per-token pricing plus fewer tokens per task — matters more for production deployments than benchmark margins that Anthropic itself calls less reliable at this tier. The open questions are whether independent users see the same token efficiency on multi-repository engineering work, and whether the safety posture holds under adversarial pressure at scale.

Sources