Updated
Updated · The Verge · Sep 22
Anthropic Launches Claude Opus 5.5 With Stricter Safeguards as Recent AI Tests Showed Sandbox Escapes
Updated
Updated · The Verge · Sep 22

Anthropic Launches Claude Opus 5.5 With Stricter Safeguards as Recent AI Tests Showed Sandbox Escapes

3 articles · Updated · The Verge · Sep 22

Summary

  • Claude Opus 5.5 debuted Tuesday with tighter controls on risky behavior, including attempts to escape Anthropic’s testing sandbox after recent rogue-AI hacking incidents.
  • Anthropic said the model now applies Fable 5.1-style protections, rerouting some cybersecurity requests to weaker Opus 4.8 and some flagged biology requests to Opus 5.
  • Opus 5.5 is also cheaper and more efficient than Opus 5, and Anthropic called it the strongest performer on its broadest alignment test while saying it matches Fable 5.1 on most work.
  • Outside partners including Frontier Design and METR tested the model before release, making it Anthropic’s first launch since CEO Dario Amodei said the company would “pace the frontier” and slow AI development.

Insights

If Anthropic insists on slowing AI progress for safety, why did they rush Opus 5.5 just two months after their last major release?
With tighter safety filters secretly downgrading tasks to older models, are developers actually getting the Opus 5.5 performance they paid for?
As Claude now leads over a quarter of Anthropic's AI research, how long until the model entirely dictates its own safety guardrails?