Updated
Updated · MIT Technology Review · Aug 26
AI Models Fail 3D, Logic and ARC Tests Despite Near-Perfect Gains on 1 Puzzle
Updated
Updated · MIT Technology Review · Aug 26

AI Models Fail 3D, Logic and ARC Tests Despite Near-Perfect Gains on 1 Puzzle

1 articles · Updated · MIT Technology Review · Aug 26

Summary

  • Spatial-reasoning tasks remain a major weak spot: language models still fail badly at mental-rotation puzzles that humans routinely solve by manipulating 3D objects.
  • 2024 and 2025 studies also show models stumble on slight twists to familiar logic problems, often recalling memorized patterns instead of adapting to new details in Knights-and-Knaves and SimpleBench tests.
  • ARC-AGI highlights a similar gap in abstract visual reasoning: models perform better when grids are encoded as numbers rather than images, suggesting they still struggle with basic visual concepts.
  • Scale adds another limit. Apple researchers found models could handle simple Tower of Hanoi and river-crossing tasks, but performance deteriorated once puzzles involved 6 or more disks or people.
  • The broader picture is mixed: some models improved from solving just 18% of New York Times Connections puzzles in late 2024 to near-perfect results by early 2025, yet humans still outperform them on several reasoning benchmarks.

Insights

As AI masters standard benchmarks, what hidden cognitive blind spots are we missing by relying on public leaderboards instead of true reasoning tests?
Why should we trust AI with real-world decisions when slight wording changes in a basic logic puzzle completely break its reasoning?
If AI perfectly solves complex word games but fails a simple river-crossing riddle, are we mistaking memorization for actual intelligence?