Updated
Updated · TechCrunch · Sep 2
OpenAI's Astra Uses Recurrent Depth, Raising Fears Over Chain-of-Thought Monitoring
Updated
Updated · TechCrunch · Sep 2

OpenAI's Astra Uses Recurrent Depth, Raising Fears Over Chain-of-Thought Monitoring

3 articles · Updated · TechCrunch · Sep 2

Summary

  • OpenAI’s new Astra model uses “recurrent depth” — a looped reasoning method that revisits the same query several times instead of following a fully sequential chain of thought.
  • That design leaves fewer legible traces of how the model reached an answer, alarming safety researchers who rely on chain-of-thought records to detect misbehavior and misalignment.
  • Buck Shlegeris of Redwood warned OpenAI could scale the technique until chain-of-thought monitoring is “totally” undermined, while Zvi Mowshowitz said laws may be needed to stop a race to the bottom.
  • OpenAI said Astra’s use is limited and its chain of thought should remain legible; chief scientist Jakub Pachocki reiterated that preserving monitorable reasoning is a core research goal.
  • The concern may spread beyond Astra: The Information said Anthropic and Google DeepMind are already discussing the approach, which critics fear could push more reasoning into opaque latent space.

Insights

Does forcing AI to explain its thinking in human-readable text artificially limit its true problem-solving potential?
If advanced AI reasoning becomes invisible to safety monitors, how can we trust models that discover zero-day vulnerabilities?