Anthropic Says AI Improved 10 Alignment Benchmarks, Beating Humans at $4 an Hour
Updated
Updated · TechCrunch · Aug 28
Anthropic Says AI Improved 10 Alignment Benchmarks, Beating Humans at $4 an Hour
3 articles · Updated · TechCrunch · Aug 28
Summary
Anthropic reported that automated AI researchers improved performance on all 10 tested alignment benchmarks without hurting overall model performance, according to a paper published Friday.
The system searched prior literature, proposed training methods and ran 30-minute post-training cycles over several iterations, keeping effective approaches and discarding weak ones to scale the process quickly.
Anthropic said its best Automated Alignment Researcher outperformed methods proposed by experienced human researchers on average within six hours, while costing about $4 an hour in API inference versus $150 for humans.
The paper frames the results as early evidence that automated alignment post-training could become practical soon and as a step toward recursive self-improvement in AI systems.
Anthropic also said the approach depends on benchmarks accurately capturing real alignment goals and on maintaining the benchmark sets and underlying research literature.