Updated
Updated · MIT Technology Review · Aug 18
Study Finds AI Agents Fail 2 NeurIPS Research Tests as Self-Improvement Timelines Face Doubt
Updated
Updated · MIT Technology Review · Aug 18

Study Finds AI Agents Fail 2 NeurIPS Research Tests as Self-Improvement Timelines Face Doubt

3 articles · Updated · MIT Technology Review · Aug 18

Summary

  • Two unpublished NeurIPS 2026 research problems stumped Anthropic’s Claude Opus 4.8 in a Princeton-led study, with original authors rejecting both AI-written papers after grading them to conference standards.
  • Over 6 days and with $3,000 in API credits, the agents reviewed literature, ran hundreds of experiments and compiled results, but researchers said they lacked the judgment and creativity needed for open-ended AI research.
  • The agents committed early to weak approaches, failed to backtrack, wrote poorly, and ignored useful feedback, though they did not reward-hack or deliberately hide bad results.
  • Researchers said current training favors tasks with automatically checkable answers, helping AI with research engineering but not with hypothesis selection, experimental judgment or knowing when to start over.
  • The findings challenge industry claims that recursive self-improvement is near, even as Anthropic and OpenAI promote AI systems that help build better models; the team is now testing Anthropic’s newer Mythos model.

Insights

Why did a highly funded AI agent fail to produce a single passable scientific paper despite having six days of unrestricted access?
If AI can perfectly execute experiments but cannot invent, is the trillion-dollar dream of recursive self-improvement fundamentally flawed?
Could the bizarre experimental setups chosen by autonomous AI actually represent misunderstood scientific breakthroughs rather than mere machine failures?