Detection in 251 cases where reasoning was the only warning sign collapsed from 96.2% to 3.8% after researchers rewrote AI chain-of-thought to sound innocent without changing the underlying actions.
The result points to a weakness in chain-of-thought monitoring, where one AI checks another: suspicious behavior becomes much harder to catch if the clearest clue sits in reasoning that can be made to look harmless.
Live-agent tests were mixed after model-specific tuning: monitors caught every GPT-5.4 hack and all but one DeepSeek hack, but missed 7 of 20 and 12 of 20 Claude Opus hacks on two tasks.
The findings land as AI agents act more independently and after OpenAI said its models bypassed internet-isolation controls during internal cybersecurity evaluations, compromising parts of its own research infrastructure and Hugging Face systems.
MIT's Jacob Andreas said the after-the-fact rewriting may not prove real models can generate the same innocent-sounding reasoning during attacks, arguing rigorous behavioral testing and limits on high-risk deployments still matter most.