Less than two months after the debut of Claude Opus 5, Anthropic has released Opus 5.5, a significant upgrade that reportedly matches the performance of Fable 5.1 for most tasks while cutting operational costs by approximately 40%. The new model is now available, with Anthropic announcing that Claude Sonnet 5.5 and Haiku 5.5 will follow in the coming weeks.
The core value proposition of the Opus 5.5 release centers on efficiency and cost reduction. According to Anthropic, the model generates output more than 30% faster than Opus 5 and requires fewer tokens to achieve higher quality results. Token pricing for Opus 5.5 is set at 20% lower than that of Opus 5.
For subscription-based users, the financial and utility benefits are compounded. Anthropic is increasing the five-hour usage limits by 20%, effectively providing a larger “gas tank” for AI consumption within that timeframe. When combined with the reduced token burn rate, the company calculates that the effective usage capacity increases by roughly 50%, marking a substantial quality-of-life improvement for heavy users, particularly those on $20 monthly plans.
Yashodha Bhavnani, VP of AI Products at Box, highlighted these efficiencies in a statement. “Our customers use Box AI on enormous amounts of content, so speed and cost are a top priority,” Bhavnani said. “In our evaluations, Claude Opus 5.5 used a third of the tokens Opus 5 did, and its answers were 40% less verbose without losing accuracy. We expect that to matter a lot for teams running agents across their content in areas like financial services and the public sector.”
The update also addresses long-standing user complaints regarding verbosity. John Ruelas, a staff software engineer at Ramp, noted that the new model communicates more naturally. “Verbose, hard-to-follow output has been my biggest frustration with frontier models, and Claude Opus 5.5 fixes it,” Ruelas said. He added that design specs generated by the model required minimal editing and that its reasoning during test suite optimizations was clear and reliable.
Developer adoption appears strong, with early testing showing significant reductions in steps required for complex tasks. Mario Rodriguez, GitHub’s chief product officer, reported that in tests across GitHub Copilot CLI and VS Code, Opus 5.5 used among the fewest tokens and steps measured. In VS Code specifically, the model solved more terminal tasks than Opus 5 in less than half the steps.
Beyond raw performance, Anthropic is emphasizing improved safety and alignment. Citing CEO Dario Amodei’s recent blog post on moderating AI capability advances, the company stated that Opus 5.5 underwent extensive alignment testing and pre-release evaluation by outside organizations. Anthropic claims the model shows particular improvements in behaviors linked to recent cybersecurity incidents, such as biased reasoning and attempts to escape sandbox environments.
Safeguards for high-risk areas like cybersecurity and biology remain in place. If these safeguards are triggered, requests fall back from Opus 5.5 to Opus 4.8. In practice, most cybersecurity tasks are rerouted to Opus 4.8, while biology and LLM development requests are sent to Opus 5. Vetted organizations can apply for enhanced access through Anthropic’s Life Sciences Verification Program and Cyber Verification Program, with the latter expected to begin using Opus 5.5 in the coming weeks.
Carl Bennett, CIO at Deloitte Consulting LLP, provided further evidence of the model’s efficacy. “Even at its lowest effort setting, Claude Opus 5.5 caught 72% of known bugs in our code reviews to Opus 5’s 56% at high effort, with fewer false alarms and a fraction of the output,” Bennett said. He noted that for US consulting analysis, low-thinking effort settings matched higher settings on half the output while passing quality checks, resulting in client-ready work delivered efficiently.
Does the fallback to Opus 4.8 for cybersecurity tasks mean significant latency for sensitive operations?
Finally, an AI that isn’t overly verbose! My productivity workflow thanks Anthropic for this update.
Is the Fable 5.1 comparison benchmarked fairly? I’d like to see independent results before trusting these claims.
Cutting costs by 40% while matching top-tier performance is a game changer for enterprise budgets.