Anthropic Researcher Says LLMs Build World Models, Planning 1 Token at a Time
Updated
Updated · O'Reilly Media · Sep 16
Anthropic Researcher Says LLMs Build World Models, Planning 1 Token at a Time
1 articles · Updated · O'Reilly Media · Sep 16
Summary
Emmanuel Ameisen said Anthropic’s interpretability work shows large language models rely on readable internal world models, with activations revealing concepts and plans during token-by-token generation.
Examples included Claude inferring San Francisco from the “GG bridge,” aiming at a rhyme before finishing a couplet, and showing abstract features for concepts like the Golden Gate Bridge across text, languages and images.
Ameisen said researchers can intervene in those features and change outputs, while hidden activations also flagged fake search results and deceptive behavior even when the model’s visible response did not.
He argued the findings challenge the idea that LLMs are just next-token predictors or fuzzy lookup tables, while also raising questions about which parts of human cognition are similarly mechanistic.
The discussion suggested AI may already handle substantial coding better than six months ago, but still struggles more with open-ended research—making problem selection and curiosity a possible human edge.