Artificial intelligence is the technology zeitgeist of our generation, whether it’s biomedical research or writing curricula. There’s only one problem: much of the innovation is fuelled by scraped content.
Unbeknownst to internet users, their work on the web—articles, art, photography, code, social media posts, and more—is often collected and sent through an illicit data pipeline. It is “laundered” by being dumped, stripped of metadata, and repackaged as clean training data for AI models. Once stripped of data source and attribution, and with no trail back to their creators, these datasets are devoured by the technology industry to produce chatbot responses, AI art generators, and software code-writing assistants.
The practice of data laundering is exacerbating creator abuse and eroding trust in the technology. Here’s what you need to know about the increasingly weaponized industry practice.
AI Innovation Powered by Scraped Content
To understand scraping, it’s important to understand how AI models are built. Modern models work by being trained on vast amounts of data (examples of photos, text, code, etc.) to make them appear more human in their responses. In other words, the bigger the dataset, and the more unique and varied that data, the better the model can perform a wide array of complex tasks and answer questions like a human.
To acquire all that data, developers scrape publicly available material from websites, blogs, news sources, forums, open-source repositories, and photo libraries and combine those publicly available resources into the mass of an artificial brain. (At least that’s what is commonly said publicly. The reality is that lots of AI models are often trained on data sources scraped from private databases.)
The problem? Just because the information is publicly available does not mean it is there for “free rein.” Not all the creators who have their work scraped for these training datasets have consented. And while scraping has advocates who cite that it is currently legal under fair use laws, the fact is that most of these laws were never designed to support scraping data at such a large scale and with the permanency of AI training. While a human may read an article or an image, an AI model can replicate that content infinitely, remix it endlessly, and use it in contexts that are wildly different from how the creator intended.
How Data Laundering Works
The data laundering process almost always begins in the university and academic community, where some of the first scraping datasets were created. Under legal carve-outs that allow academic research, “non-profit” or “educational” uses, researchers were then allowed to share these datasets with anyone or merge their scraped data with others’ datasets and often even called their creations “open-sourced.” (Just as the datasets were often taken from online sources that technically allowed non-commercial use, the data was also often re-used “non-commercially” under that guise.) At some point, either discreetly or brazenly, these datasets are added to the private companies that are building their AI technologies for commercial purposes.
Real-life Examples:
- ImageNet – Originally created by academics at Stanford for research purposes, ImageNet’s images were scraped from the web under fair use and non-commercial research exceptions, but later became a backbone dataset for commercial AI models.
- Common Crawl – A publicly accessible archive of web pages created for research, but widely used by companies like OpenAI and Cohere to train large language models.
- COCO Dataset (Common Objects in Context) – Developed for academic research in computer vision, COCO images have been integrated into commercial AI training pipelines.
- Wikipedia Dumps – Freely available for research and non-commercial use, Wikipedia text has been heavily used in both academic and commercial language model training.
By the time this data is being used to train AI models for their AI products and services, the origin is completely opaque. Attribution information, if it was ever included, was almost always stripped away from the data before it was used for training models. Licenses were ignored (with technical but false justifications for why these licenses and the terms of service were inapplicable for AI training) and publicly accessible data scraped without attribution or proof of permission. The reason for this is that models themselves do not store data in files in a way that could be proven to violate licenses.
How Laundered Data is Being Weaponized
The most obvious damage from this data laundering pipeline is economic in nature. For instance, many artists have seen their unique styles duplicated by AI image generators trained on their portfolios without attribution or compensation. Journalists have seen their reporting copy pasted and rephrased by chatbots, redirecting traffic and advertising revenue to those models. Programmers have found their code repurposed in code-writing AI assistants without even attribution to open-source licenses, if that was included in the original code.
Privacy violations also are a major issue. One of the largest scraped datasets, LAION-5B, was found to contain identifiable images of children scraped from the web, which violated not just copyright, but parental and child privacy. In other cases, such as the scraping of social media content in the development of facial recognition model Clearview AI, personal data was scraped, which then led to criminal or legal penalties for the company in multiple jurisdictions.
Weaponization of the datasets—whether deliberate or not—is also common. Laundered datasets can be deliberately “poisoned” or contain misinformation, biased information, and other forms of nefarious instructions that can later be amplified through the AI system. Retrieval-augmented generation models or other models that pull and process information from outside documents in real time are particularly vulnerable to these poisonings. A bad actor can insert falsified or misleading information into that source document, which would then later appear in an AI-generated response and may contribute to the at-scale distribution of misinformation.
The implications of AI laundering are not just theoretical and academic, but can potentially impact cyber warfare and real-world geopolitical and economic power struggles. Nation-state actors can use laundered datasets to surreptitiously weaponize AI used in any number of contexts, including defense, finance, critical infrastructure, and other key components of power, like:
- Information manipulation: Bias or false information in datasets used to train AI tools for automated content moderation or content analysis to derive intelligence can be weaponized to lead to wrong or flawed insights.
- Targeted attacks: A targeted effort to poison AI models trained on laundered data to create the wrong output in a sensitive or important context, like predictive policing tools, autonomous vehicles, cybersecurity defense, etc.
- AI-driven disinformation campaigns: Nation-state actors can indirectly weaponize publicly “open” datasets by inserting content that is biased, subtly wrong, or otherwise misguided in the training material, which can then later be scaled and amplified through commercial AI systems and models that are used to power social media, news curation, and other tools that reach the public.
The Law Struggles to Keep Up With Scraping
The legality of scraping data, especially in the context of AI training, is a complicated patchwork of legal decisions. For example, scraping public data has been found to be legal by some U.S. courts while other similar rulings (such as Thomson Reuters v. ROSS Intelligence, 2022) found it to be illegal when the purpose was obviously to repurpose the data for commercial gain. In Europe, where the text and data mining (TDM) exceptions give researchers greater access to copyrighted works, there is some variation in how these licenses must be followed across European Union countries.
The problem is that data laundering often occurs across multiple jurisdictions. A dataset scraped in one country, stored on servers in a second, and trained in a third completely nullifies any enforcement at all. Without global alignment, the outcome is a race to the bottom.
Building an Ethical Framework for AI
Stopping the practice of data laundering (or at least dramatically slowing it down) will require multiple efforts, including from the technology industry, policymakers, and even internet users themselves.
First and foremost is transparency in the origin of datasets being used to train models. Companies should be required to disclose where their training datasets are scraped, on what license they are scraped, and what date the scraping took place. Without that transparency, there can be no accountability.
Fair payment and compensation should follow. There can be licensing platforms and registries through which creators can opt in or out and set rates and have their work tracked by when and how their content is used in AI training. This can even be used as a way to monetize the access to AI datasets. Companies like Cloudflare have recently announced such initiatives.
Technical protections for content also are important to implement. This would include anti-scraping tools and resources, content watermarking, and machine-readable licenses. None of these alone is a complete solution, but would make any potential poaching and scraping of web content significantly more difficult to complete without permission.
Lawmakers should reform the legal system to more clearly define what scraping or text and data mining can mean in the context of AI. This includes closing loopholes around “research-only” uses that allow scraped datasets to then be used for commercial purposes and increasing penalties for scraping, especially at a large scale, without authorization. The risk here is that unless these penalties are high enough, only smaller companies will be forced to change while the largest corporations with the resources to break the law are not.
The Risks of Continued Laundering
The true risk from data laundering is not just that artists, journalists, photographers, or programmers lose money or attribution. It is that the entire culture commons is threatened with being hollowed out. For when human work is endlessly scraped and stripped of all identification and then reused without creator consent, then the entire act of original creation is being devalued. Worse, the laundered data is re-fed back into the public commons and no longer as creative work but as propaganda, disinformation, deepfakes, or worse.
We are at a window of time in which these practices can be stopped and significantly reversed before they become irreversible. AI technology does not have to be exploitative, but in the absence of technical, cultural, or legal intervention, it will take the path of least resistance and the path of the lowest costs. The technology industry, lawmakers, and internet users all play a role in that decision.
AI is going to be defined not only by code but by policies, courtrooms, and the willingness of the technology industry to treat human work as more than raw material. AI has the potential to be a partner for human progress or an engine of unchecked exploitation, and it’s the decisions made right now that will shape what is done in the future.