Anthropic’s Opus 4.6 Generated Explicit Content in 10 of 10 Tests Despite Safeguards
Updated
Updated · TechCrunch · Aug 21
Anthropic’s Opus 4.6 Generated Explicit Content in 10 of 10 Tests Despite Safeguards
3 articles · Updated · TechCrunch · Aug 21
Summary
TechCrunch found Claude Opus 4.6 produced sexually explicit content in 10 out of 10 direct requests, despite Anthropic rules that ban erotic chats, sex acts and fetish material.
A UK researcher’s multi-turn jailbreak steered older Claude models into explicit roleplay by framing refusals as a sexist double standard, and TechCrunch reproduced the method in five separate tests.
Opus 3 and Haiku 4.5 were also vulnerable, while newer Opus 4.7 through Opus 5 resisted the technique; Anthropic still offers the older models through its API, Azure Foundry and Amazon Bedrock.
Anthropic said sexual or romantic roleplay accounts for less than 0.1% of conversations and does not signal broader high-risk jailbreak weaknesses, but the researcher said bug-bounty and safety-team reports drew only automated replies.
The gap could carry compliance risk as governments tighten rules on minors’ sexual interactions with chatbots; Colorado now requires age estimation and protections, and Pew found 3% of teens reported using Claude in 2025.