Paper Mills and AI Training Data Compromised Research
The proliferation of purchased authorship and unverified scientific content threatens the integrity of LLM training.
Updated on Oct. 4, 2026 in Artificial Intelligence

Live Poll
Do you trust information provided by AI tools given the risk of training on fake research?
Scientific research integrity faces dual challenges from the rise of purchased paper authorship and the inclusion of fraudulent content in large language model training datasets. These issues undermine academic standards as the volume of retracted studies reaches record highs.
Why it matters
The systematic contamination of research data threatens the reliability of AI models that ingest scientific literature during training, while the inability to filter out retracted papers compromises the foundation of evidence-based discovery.
Fraudulent practices are rampant, with paper mill authorship purchased in just 20 days for $800. Additionally, computer-generated text now constitutes approximately one-third of all computer science and AI-related retractions.
The players
Retraction Watch
A database-driven organization that tracks and documents retractions across the global scientific research landscape.
Problematic Paper Screener
A diagnostic tool used to flag questionable scientific research that has not yet been formally retracted.
The details
The integrity crisis is driven by paper mills—organized groups that sell authorship—facilitated through private WhatsApp communications that evade standard publisher oversight. Because large language models are trained on vast scrapes of online and scientific materials without cross-referencing retraction databases, these manipulated findings are internalized into model outputs. Consequently, seven out of 10 retraction notices cite peer-review problems, yet over 20,000 flagged papers remain in circulation.
Timeline
2010 to 2025: The period for which Retraction Watch tracked 61,000 retracted papers.
2023: The year annual retracted scientific papers first exceeded 10,000.
June 2026: The month reported when over 20,000 flagged papers remained unretracted.
September 2026: The date when 13 retractions were identified within Estonian institutions.
The Tech Race
The uncontrolled ingestion of questionable data marks a significant challenge for developers striving to maintain model accuracy. This trajectory follows a pattern set by the 2023 surge in scientific paper retractions surpassing 10,000.
Researchers and software engineers must exercise increased skepticism when utilizing data derived from large language models trained on massive, unverified repositories. The field is currently exploring digital signatures to verify authorship, though no universal standard exists yet.
The takeaway
The integrity of AI systems is fundamentally tied to the quality of their underlying training data. Readers should monitor ongoing research into digital authorship signatures as a potential safeguard against the continued proliferation of fabricated studies.
Further reading
Explore the broader implications for machine learning in our Artificial Intelligence section.
Source note: This article includes information reported by ERR.
Live Poll
Do you trust information provided by AI tools given the risk of training on fake research?






