Harvard and Chutes Released Massive AI Usage Dataset
A new dataset tracking 6.1 billion requests reveals that most AI user behavior repeats within 15 minutes.
Updated on Sept. 21, 2026 in Artificial Intelligence

Live Poll
Are you using automated AI systems to perform specific, targeted tasks more frequently?
Harvard University and Chutes have released a comprehensive dataset containing 6.12 billion LLM requests spanning over 9,000 models. This research-grade collection covers data gathered between April 2025 and April 2026.
Why it matters
The dataset provides an unprecedented look at how users interact with AI models at scale, highlighting a specific trend in short-term repetitive querying. This information helps developers optimize latency and caching strategies for future AI infrastructure.
The dataset accounts for 35.8 trillion input tokens and 2.52 trillion output tokens across 9,174 models. Performance benchmarks indicate that 99% of repeat requests occur within a 15-minute window, while average response lengths shrunk from hundreds of tokens to fewer than 100.
The players
Harvard University
An Ivy League institution with an extensive research portfolio in computer science and data privacy.
Chutes
An infrastructure firm specializing in AI request routing and usage analytics.
Bittensor Subnet 64
A decentralized machine learning network providing the infrastructure for tracking these large-scale model requests.
The details
The data was collected by incentivizing users with a 25% discount for contributing usage logs during an opt-in period. To preserve privacy, the researchers implemented a system where user identifiers rotate every three months. This allows for longitudinal analysis of request patterns while limiting the ability to track individual user behavior over long periods.
Timeline
Dataset collection began on April 11, 2025.
Data contribution opt-in occurred from March to July 2026.
Dataset collection concluded on April 12, 2026.
The full dataset was released on September 21, 2026.
The Tech Race
This release follows the open-access protocols established by the Harvard University research data repository. It marks a significant departure from proprietary model-provider logs by offering third-party researchers access to multi-model usage behavior.
Developers and researchers can access the full dataset via GitHub or the Harvard S3 bucket to refine AI application performance. The findings suggest that caching systems for AI should prioritize extremely short-term request windows to handle the 99% of repeat queries identified.
The takeaway
The high volume of repeat requests within short 15-minute windows suggests that AI infrastructure is currently over-taxed by redundant processing. Developers should monitor future updates to this dataset to see if response lengths continue to trend downward as efficiency improves.
Further reading
For more on the current state of model research, visit our Artificial Intelligence section.
Live Poll
Are you using automated AI systems to perform specific, targeted tasks more frequently?









