ua en ru

Centuries-old printing error now poisons AI systems, study finds

Thu, September 10, 2026 - 00:40
3 min
How did a simple error confuse scientists?
Centuries-old printing error now poisons AI systems, study finds A meaningless phrase from the past undermines science (photo: Magnific)

Scientists have discovered dozens of peer-reviewed papers containing a completely fabricated term that does not exist in scientific discourse. It originated from errors in optical character recognition and machine translation and later entered databases used to train AI models, according to Futura.

Phantom term: from a 1950s typo to Iran

It all began in the 1950s, when two articles from the Bacteriological Reviews journal were digitized.

Due to an error in optical character recognition (OCR), the word "vegetative" from one column was accidentally merged with the word "electron" from the adjacent column. This led to the creation of a fictional term that does not exist in real scientific language.

For decades, the phantom went unnoticed until it surfaced in English-language abstracts of Iranian scientific publications in 2017 and 2019.

The cause was a translation error: in Persian (Farsi), the words "vegetative" and "scanning" differ by only one diacritic mark. This small mistake introduced the false phrase into international scientific circulation.

"Infection" of AI models

According to Google Scholar, this nonsensical term appeared in at least 22 scientific publications.

Researchers decided to check whether language models had absorbed this error.

During testing with original article excerpts, GPT-3 automatically completed sentences with "vegetative electron microscopy", while older systems like GPT-2 or BERT did not show the same tendency.

The anomaly has also persisted in newer versions - GPT-4o and Anthropic's Claude 3.5.

The main source of this "infection" was the massive open web dataset Common Crawl, used to train most modern neural networks.

Why are "digital fossils" dangerous for science?

This case shows how vulnerable AI algorithms are to systematic errors from the past.

Modern automated text verification tools now flag the phrase "vegetative electron microscopy" as a marker of AI-generated content.

However, such systems are only capable of detecting errors that have already been identified by humans.

To preserve the integrity of scientific work, experts urge AI developers to open up their training datasets and scientific publishers to strengthen peer review to prevent the self-replication and entrenchment of fake data in digital space.

Or read us wherever it's convenient for you!