AI Labs Faced Data Scarcity as Scraping Friction Rose
As high-quality training sources tighten access, AI labs have pivoted toward aggressive licensing and data acquisition.
Updated on Oct. 7, 2026 in Artificial Intelligence

Live Poll
Should AI companies be required to obtain explicit consent before training on your public content?
AI laboratories have relied on extensive web scraping to build large language models, but publishers are increasingly restricting access via paywalls and robots.txt protocols. This shift follows reports that some news outlets saw click-through rates fall by 51 to 94 percent after AI product launches.
Why it matters
The scarcity of high-quality, publicly available data is forcing AI labs to secure content through private, multimillion-dollar licensing deals. This transition marks a move away from open-web scraping toward a pay-to-play model for training material.
AI training sets have frequently utilized broad scraping; for example, one Common Crawl dataset included over 2 million documents from nytimes.com, while another mid-training set contained 91,000 works from the Times, Daily News, and the Center for Investigative Reporting.
The players
Brent Hecht
Microsoft Director of Applied Science who has characterized current AI training practices as theft.
A social media platform that currently licenses its historical user data to companies like Google and OpenAI for model training.
Shutterstock
A stock photography provider that has transitioned its business model to include significant data licensing revenue for AI training.
Dataset Providers Alliance
An organization that advocates for an opt-in system regarding the use of web content for AI model training.
The details
AI labs typically employ web scrapers to aggregate vast quantities of text for model training, often ignoring robots.txt files—a standard for indicating which parts of a website should not be accessed by automated bots. To mitigate the resulting data drought, companies are now negotiating private licensing agreements or purchasing entire archives from defunct businesses. The Dataset Providers Alliance is currently advocating for an industry-wide shift toward an opt-in system for all training data acquisition.
Timeline
2024: Shutterstock reported $138 million in data licensing revenue.
Summer 2026: The Dataset Providers Alliance formed.
October 2026: Gizmodo coined the term slurp juice to describe aggressive scraping.
The Tech Race
The AI training dataset market is projected to grow to $4 billion as companies compete for high-quality data. This race for exclusive information marks a departure from the early, unconstrained era of web scraping.
Readers may notice increased paywalls and more restrictive terms of service on their favorite news websites as publishers move to block AI scrapers. These measures aim to protect content, though they may limit free access to high-quality information online.
The takeaway
The move toward an opt-in data ecosystem could fundamentally alter the economics of large language model development. Watch for the growth of the AI training dataset market, which is expected to multiply significantly by the early 2030s.
Further reading
For more on how data acquisition is shifting, see the latest updates in Artificial Intelligence.
Source note: This article includes information reported by WebProNews.
Live Poll
Should AI companies be required to obtain explicit consent before training on your public content?









