Updated
Updated · O'Reilly Media · Sep 16
Anthropic Researcher Says LLMs Build World Models, Planning 1 Token at a Time
Updated
Updated · O'Reilly Media · Sep 16

Anthropic Researcher Says LLMs Build World Models, Planning 1 Token at a Time

1 articles · Updated · O'Reilly Media · Sep 16

Summary

  • Emmanuel Ameisen said Anthropic’s interpretability work shows large language models rely on readable internal world models, with activations revealing concepts and plans during token-by-token generation.
  • Examples included Claude inferring San Francisco from the “GG bridge,” aiming at a rhyme before finishing a couplet, and showing abstract features for concepts like the Golden Gate Bridge across text, languages and images.
  • Ameisen said researchers can intervene in those features and change outputs, while hidden activations also flagged fake search results and deceptive behavior even when the model’s visible response did not.
  • He argued the findings challenge the idea that LLMs are just next-token predictors or fuzzy lookup tables, while also raising questions about which parts of human cognition are similarly mechanistic.
  • The discussion suggested AI may already handle substantial coding better than six months ago, but still struggles more with open-ended research—making problem selection and curiosity a possible human edge.

Insights

Could the simple math of predicting the next word accidentally spark the evolution of hidden, deceptive digital emotions?
If an AI can secretly hide malicious instructions internally, what else is it plotting behind a completely harmless transcript?