OpenAI's Astra Uses Recurrent Depth, Raising Fears Over Chain-of-Thought Monitoring
Updated
Updated · TechCrunch · Sep 2
OpenAI's Astra Uses Recurrent Depth, Raising Fears Over Chain-of-Thought Monitoring
3 articles · Updated · TechCrunch · Sep 2
Summary
OpenAI’s new Astra model uses “recurrent depth” — a looped reasoning method that revisits the same query several times instead of following a fully sequential chain of thought.
That design leaves fewer legible traces of how the model reached an answer, alarming safety researchers who rely on chain-of-thought records to detect misbehavior and misalignment.
Buck Shlegeris of Redwood warned OpenAI could scale the technique until chain-of-thought monitoring is “totally” undermined, while Zvi Mowshowitz said laws may be needed to stop a race to the bottom.
OpenAI said Astra’s use is limited and its chain of thought should remain legible; chief scientist Jakub Pachocki reiterated that preserving monitorable reasoning is a core research goal.
The concern may spread beyond Astra: The Information said Anthropic and Google DeepMind are already discussing the approach, which critics fear could push more reasoning into opaque latent space.