Nearly 3 in 10 ARC-AGI-1 puzzles were solved by experimental model BDH-CQ in two attempts, with the system reasoning internally instead of generating chain-of-thought text.
A fixed-size memory lets BDH-CQ absorb example puzzles without revisiting all prior context, aiming to cut the token-heavy computation that makes many AI reasoning systems slower and costlier.
At about $0.00070 per puzzle, the researchers estimate BDH-CQ runs at roughly one-eleventh the cost of GPT-5.6 Luna on the same benchmark, though the cost calculations were not directly comparable.
ARC-style tests ask models to infer visual rules from examples; BDH-CQ handled some shape rotations and movements better than color changes, ordering tasks and nested rule combinations.
Researchers and outside experts called the result promising but preliminary: the paper is unreviewed, the model is specialized for ARC problems, and hidden reasoning makes its decisions harder to inspect.