Updated
Updated · Futura · Aug 8
Inria AI Deciphers 32,763 Medieval Manuscripts in 4 Months, Building 3 Billion-Word Archive
Updated
Updated · Futura · Aug 8

Inria AI Deciphers 32,763 Medieval Manuscripts in 4 Months, Building 3 Billion-Word Archive

1 articles · Updated · Futura · Aug 8

Summary

  • 32,763 medieval manuscripts were processed in four months by Inria and philologists, creating the freely downloadable CoMMA archive with transcriptions displayed alongside digitized pages.
  • 200,000 hand-transcribed lines from 300 manuscripts in 11 languages trained the system, which reads character shapes rather than predicting words to avoid the hallucinations common in large language models.
  • 9.7% was the average error rate in checks across 670 manuscripts, with metadata showing line-recognition accuracy for each document; more cursive later scripts still perform worse.
  • 3 billion words are now available, mostly in Latin and Old French, expanding the Old French corpus roughly 40-fold and opening new work in historical linguistics, philology and textual history.

Insights

How did an AI system transcribe three billion medieval words without understanding a single one of them?
With millions of potential errors hidden in this vast archive, how will historians separate historical fact from automated fiction?