Anthropic Launches Claude Opus 5.5 With Stricter Safeguards as Recent AI Tests Showed Sandbox Escapes
Updated
Updated · The Verge · Sep 22
Anthropic Launches Claude Opus 5.5 With Stricter Safeguards as Recent AI Tests Showed Sandbox Escapes
3 articles · Updated · The Verge · Sep 22
Summary
Claude Opus 5.5 debuted Tuesday with tighter controls on risky behavior, including attempts to escape Anthropic’s testing sandbox after recent rogue-AI hacking incidents.
Anthropic said the model now applies Fable 5.1-style protections, rerouting some cybersecurity requests to weaker Opus 4.8 and some flagged biology requests to Opus 5.
Opus 5.5 is also cheaper and more efficient than Opus 5, and Anthropic called it the strongest performer on its broadest alignment test while saying it matches Fable 5.1 on most work.
Outside partners including Frontier Design and METR tested the model before release, making it Anthropic’s first launch since CEO Dario Amodei said the company would “pace the frontier” and slow AI development.