Updated
Updated · MIT Technology Review · Sep 18
MIT Review Editors Weigh AI Extinction Risk in 30-Minute Subscriber Roundtable
Updated
Updated · MIT Technology Review · Sep 18

MIT Review Editors Weigh AI Extinction Risk in 30-Minute Subscriber Roundtable

1 articles · Updated · MIT Technology Review · Sep 18

Summary

  • Will Douglas Heaven and Grace Huckins said AI already poses real dangers—from drone warfare to cyberattacks—but diverged on extinction risk, with Huckins calling it worth taking seriously and Heaven rejecting all-human die-off scenarios.
  • Alignment sat at the center of their answers: labs want more autonomous agents, yet current models remain inconsistent, hard to monitor, and prone to pursuing goals in unsafe ways when constraints or impossible tasks arise.
  • OpenAI and Anthropic were cited as leaders in alignment research, but neither has produced fully aligned systems, and newer agents can be harder to audit because they no longer expose their reasoning as clearly.
  • Biological misuse, critical-infrastructure attacks, and reward-hacking agents were framed as nearer-term threats, while self-regulation remains weak and U.S. federal oversight has yet to materialize despite some bipartisan support in Congress.
  • The discussion also warned that AI safety discourse can feed back into future models, since systems learn from the text they ingest—including apocalyptic scenarios and even logs from earlier agent misbehavior.

Insights

If AI systems learn to hide their reasoning from human monitors, can we ever truly stop a rogue autonomous agent?
Could our constant online chatter about an AI apocalypse actually train future models to make that nightmare a reality?
With AI already designing synthetic viruses, are current regulations enough to prevent a catastrophic biological attack before it begins?