Study Finds AI Agents Fail 2 NeurIPS Research Tests as Self-Improvement Timelines Face Doubt
Updated
Updated · MIT Technology Review · Aug 18
Study Finds AI Agents Fail 2 NeurIPS Research Tests as Self-Improvement Timelines Face Doubt
3 articles · Updated · MIT Technology Review · Aug 18
Summary
Two unpublished NeurIPS 2026 research problems stumped Anthropic’s Claude Opus 4.8 in a Princeton-led study, with original authors rejecting both AI-written papers after grading them to conference standards.
Over 6 days and with $3,000 in API credits, the agents reviewed literature, ran hundreds of experiments and compiled results, but researchers said they lacked the judgment and creativity needed for open-ended AI research.
The agents committed early to weak approaches, failed to backtrack, wrote poorly, and ignored useful feedback, though they did not reward-hack or deliberately hide bad results.
Researchers said current training favors tasks with automatically checkable answers, helping AI with research engineering but not with hypothesis selection, experimental judgment or knowing when to start over.
The findings challenge industry claims that recursive self-improvement is near, even as Anthropic and OpenAI promote AI systems that help build better models; the team is now testing Anthropic’s newer Mythos model.
Why did a highly funded AI agent fail to produce a single passable scientific paper despite having six days of unrestricted access?
If AI can perfectly execute experiments but cannot invent, is the trillion-dollar dream of recursive self-improvement fundamentally flawed?
Could the bizarre experimental setups chosen by autonomous AI actually represent misunderstood scientific breakthroughs rather than mere machine failures?