Updated
Updated · Futura · Sep 8
AI Models Absorb Bogus Term Found in 22 Papers as OCR and Translation Errors Fossilize
Updated
Updated · Futura · Sep 8

AI Models Absorb Bogus Term Found in 22 Papers as OCR and Translation Errors Fossilize

1 articles · Updated · Futura · Sep 8

Summary

  • Researchers found the meaningless phrase “vegetative electron microscopy” embedded in AI systems after it surfaced in 22 scientific publications and multiple model outputs.
  • GPT-3 completed prompts with the phrase, while GPT-4o and Claude 3.5 still reproduced it; earlier models such as GPT-2 and BERT did not show the same tendency.
  • The term appears to trace back to a 1950s digitization error that merged words across columns, then spread further through Farsi-to-English translation confusion in 2017 and 2019 papers.
  • Springer Nature retracted affected articles, while Elsevier first defended the term before issuing corrections, underscoring uneven publisher responses as web archives like CommonCrawl preserve such errors.
  • The case highlights how “digital fossils” can harden into scientific and AI knowledge bases, while current detection tools mostly catch only mistakes already known.

Insights

If a meaningless phrase can infect global academic databases undetected, what other digital fossils are secretly shaping modern science?
How did a 1950s scanning glitch trick the world's most advanced AI models into hallucinating a fake scientific method?
Can we ever truly erase a digital mistake once it becomes permanently embedded in the training data of powerful AI systems?