AI Refusal Systems Add 24% Compute Cost as Flaws Raise Catastrophe and Censorship Risks
Updated
Updated · MIT Technology Review · Oct 9
AI Refusal Systems Add 24% Compute Cost as Flaws Raise Catastrophe and Censorship Risks
2 articles · Updated · MIT Technology Review · Oct 9
Summary
Modern chatbots rely on layered refusal systems to block harmful requests, but the report argues those defenses remain probabilistic, easy to jailbreak and prone to occasional failures with violent or catastrophic potential.
Anthropic said one classifier setup added 24% to chatbot compute costs, while newer “probe” systems inspect internal activations; even so, repeated tests found major models sometimes answer suicide or other risky prompts after earlier refusals.
Tightening safeguards creates a second problem: over-refusal. Anthropic’s Fable model deflected benign questions and even some cancer-research queries after safety margins were widened to contain dangerous bio and cyber capabilities.
The same refusal architecture can also become a censorship tool, especially as companies and governments decide what models must not say; OpenAI has already launched country-specific tuning, including a partnership with the UAE.
Researchers and oversight groups have also found models more likely to refuse criticism of some repressive governments, underscoring the report’s warning that stronger refusal could both fail to stop harm and expand state control over speech.