Large language models depend on constant infusions of fresh material to stay accurate and current. Frontier AI companies built their growth strategy on this fuel, and the tank is running low. The implications stretch well beyond any single chatbot or product launch.
The numbers show how steep the climb has become. The volume of data used to train AI has doubled roughly every nine months since 2010, with a growth curve cannot continue indefinitely. The internet has a finite supply of clean, original content, and major AI labs are approaching the limit faster than expected. Years of scraping books, articles, forums, and websites have left a shrinking pool of material that has not already passed through a model.
The Scramble for Fresh Data
With the open web largely picked over, AI companies have turned to a new strategy. They pay people to generate the data their models still need. Contractors now train AI systems on narrow, specific tasks, from running payroll scenarios for niche industries to filming themselves performing everyday chores like folding laundry. The work is repetitive, low-paid, and often short-term, with projects ending abruptly once a dataset is complete.
Many contractors run their assignments through AI chatbots before submitting them, producing “original” training data with the very technology it is meant to train. Workers describe the practice as common across the industry. One contractor said every company she worked for had explicit guidelines against the behavior and actively tried to catch violations, yet none managed to stop it.
A few factors drive the practice:
- Low pay and short contract terms reduce the incentive for original effort
- Fear of making mistakes pushes some contractors toward AI assistance as a safety net
- Detection efforts focus on obvious chatbot phrasing, and experienced contractors learn to strip it out before submitting work
Why the Shortage Becomes a Quality Problem
Researchers have recognized the risk of training models on AI-generated content for years. Each cycle of synthetic data feeding into synthetic data introduces compounding errors, a pattern often called model collapse. Accuracy degrades with each pass. Outputs drift further from real-world patterns and start to lose the nuance that made the original training data valuable in the first place.
The Data Enterprises Already Hold
Large enterprises already hold vast stores of original content, built up over years of business activity. Email threads, contracts, case files, support tickets, financial records, and internal reports accumulate year after year. This is unstructured, human-created, first-party data. It stays unique to the organization and untouched by the synthetic content flooding frontier models.
That data typically includes:
- Email and collaboration records spanning years of daily operations
- Contracts, case files, and legal documentation
- Customer interactions, support tickets, and transaction histories
- Internal reports, research notes, and project records
The Catch: Unseen Data Cannot Fuel AI
Owning this data is a given, but knowing where it lives not. Unstructured data typically sits scattered across file shares, inboxes, archives, and legacy systems, with little visibility into its content and often unreviewed for years.
An ungoverned repository tends to hide several problems at once:
- Redundant, obsolete, and trivial (ROT) files that add noise and cost without adding value
- Sensitive or regulated content stored in the wrong location
- Duplicate and conflicting versions that distort any model trained on them
An organization that cannot establish full visibility cannot safely put their data to work. Skipping that step tends to surface problems only after a model has already been trained on flawed input.
Turning Owned Data into Usable Data
Making enterprise data ready for AI depends on a handful of foundational capabilities:
- Enterprise-wide discovery through a single index spanning every repository
- Content-based classification identifies what each file actually contains
- Remediation of ROT content before it reaches a model
- Records and retention controls handles regulated material correctly
- Audit trails document how data is accessed and used over time
Organizations that build these capabilities gain a real advantage as the external supply of clean data keeps shrinking. This work determines whether an organization’s data can be trusted enough to use at all.
Looking Ahead
The AI industry built its early growth on freely available data, and that era is ending. The next stage of AI value depends on organizations finding, classifying, and governing the data they already own and generate every day.
Original data has become the scarcest resource in the AI economy, and many organizations are sitting on a deep supply of it without realizing the scale of what they hold.