Updated
Updated · MIT Technology Review · Aug 24
GPT-BERT Wins BabyLM 2024, Beating Llama 2 70B With 100 Million Words
Updated
Updated · MIT Technology Review · Aug 24

GPT-BERT Wins BabyLM 2024, Beating Llama 2 70B With 100 Million Words

1 articles · Updated · MIT Technology Review · Aug 24

Summary

  • GPT-BERT won BabyLM 2024 after pretraining on about 100 million words and still beating Meta’s Llama 2 70B on one grammar benchmark.
  • That result targets AI’s data-efficiency gap: toddlers start producing grammatical sentences after roughly 10 million to 30 million words, while frontier models train on vastly larger corpora.
  • BabyLM was built to test whether child-scale training can produce strong language models, but organizers say popular baby-like ideas such as curriculum learning have not worked as well as expected.
  • Researchers increasingly think closing the gap may require more than text—adding vision, interaction and social learning—though multimodal and interactive baby-inspired models still lag standard approaches.
  • The broader payoff could extend beyond English chatbots, helping universities and smaller-language communities build capable models with limited data while giving cognitive scientists a new tool to study human language learning.

Insights

If AI needs trillions of words to speak, do human babies possess a hidden evolutionary code that machines are missing?
Will the quest to build baby-like neural networks finally save the tech industry from its impending internet data shortage?