The quote is out
It was called ‘the largest theft of labor in human history’.
That is not a protest slogan. It is not a line from a furious newspaper editorial. It is the private description of artificial intelligence written by a Microsoft executive. The assessment, concerning the data scraping practices of its own partner OpenAI, was never intended to be seen outside the company. It is a stunning admission.
The phrase now sits in the public record. It was made public on 17 September. The words appeared in court documents unsealed as part of a major lawsuit filed by news organisations. What was once a private anxiety, typed into an internal email, is now evidence for the entire world to see. The revelation provides a brutal contrast to the carefully managed image Microsoft has cultivated, pouring billions of pounds into OpenAI and presenting its technology as the foundation of a new industrial revolution.
Microsoft’s public statements speak of empowerment. They talk of creativity. Its executives tour the world to assure customers, regulators and the public that they are building the future responsibly. They talk of safeguards. They speak of alignment. All while the company has been a primary beneficiary of the very activity one of its own people privately branded as theft. Behind the curtains, a different conversation was happening.
The language was much blunter. It was about ‘theft’. The source of this alleged theft is the very information news organisations produce, the articles and investigations they place behind paywalls to fund their work. According to the court filings, this is the material that was being used. It was being scraped. It was being fed into the machine. This private acknowledgement, coming from inside one of the two companies at the very heart of the AI boom, fundamentally reframes the debate about the technology's costs and benefits. It is no longer just critics making the accusation. The call is coming from inside the house.
They worked together
This was not a competitor’s jab. It was an ally’s fear. Microsoft and OpenAI are not rivals fighting for market share. They are locked in the technology sector’s most significant partnership, a union that has seen Microsoft pour an estimated £10.3 billion into the San Francisco based AI laboratory. This is a deep integration. Microsoft’s Azure cloud platform is the engine that powers OpenAI’s models, and in return, Microsoft gets to weave that same advanced technology, like GPT-4, directly into its own products from the Bing search engine to its Office software suite. Their fortunes are linked.
The admission’s source makes it so explosive. This is not some distant academic or a disgruntled former employee lodging a complaint from the sidelines. This observation came from within the Microsoft machine itself, the corporate behemoth that has staked its reputation and a significant portion of its future growth on OpenAI’s success. The two companies present a united front to the world. They promise a new age of productivity. Yet the court filings allege a shared, secret economy running underneath this public campaign, one built on a foundation that a key figure apparently considered illegitimate. This was a shared project.
The legal papers do not permit Microsoft to stand apart from its celebrated partner. They cannot plead ignorance. The filings allege that both companies actively participated in building vast datasets using paywalled content scraped directly from news organisations. This was a shared method. It was a mutual benefit. The very activity described internally as ‘theft’ was, according to the legal arguments, a foundational practice for a shared and immensely profitable enterprise. The goal was to build the most powerful artificial intelligence models on the planet, and the fuel for these models is data. It is vast oceans of human language and knowledge. The court filings allege that instead of paying for access to high quality, curated information, both companies simply took it. This makes the private admission of ‘theft’ an indictment not just of a partner’s methods, but of a strategy from which Microsoft itself has handsomely profited.
This is the 'doom loop'
They had a name for this cycle. A grim one. The emails called it the ‘doom loop’. This is not some abstract academic theory but a specific, catastrophic sequence of events that people inside Microsoft apparently feared was already in motion, a self fulfilling prophecy where the AI’s own voracious appetite leads directly to its eventual starvation. The loop describes a future where the technology’s success causes its own collapse. It is a self eating machine. It is a slow, technological suicide.
The process starts simply. The AI consumes information. It scrapes news articles, analyses reports, and ingests human knowledge from across the internet, all to learn how to communicate and reason. The problem, as the ‘theft’ comment makes plain, is that the filings allege it does this without paying the creators of that information, the news organisations who spend millions of pounds sending reporters to council meetings, to war zones, and to courtrooms. As users increasingly turn to AI chatbots for instant summaries instead of clicking through to the original news sites, the publishers lose the subscription revenue and advertising income that funds their entire operation. Their businesses wither. Their staff are cut. The flow of new, verified, high quality human information slows to a trickle before it stops.
This is where the loop closes on itself. The machine starves. What happens when the information ecosystem has been hollowed out, when the reliable sources have gone bankrupt and shuttered their websites. The AI needs fresh data. It needs a constant torrent of updates to remain relevant and accurate. An AI model trained on the internet of 2025 cannot tell you about the general election of 2029 or the winner of the 2030 World Cup without a steady supply of new journalism to learn from. If that journalism disappears, the AI is forced to feed on a degraded diet of lower quality user generated content, or worse, its own synthetic, recycled outputs. The model’s knowledge base stagnates and then decays. Its answers become unreliable, then circular, then simply wrong. This is the doom.
The threat is therefore existential. It affects the AI companies themselves. The internal emails show a shocking recognition that this parasitic relationship is not sustainable in the long term, that by saving money on licensing content today the companies risk destroying the very resource they need to build their products for tomorrow. A high quality information ecosystem is not an obstacle to be overcome. It is the fertile soil from which their technology grows. The Microsoft executives who discussed the 'doom loop' seem to have understood this perfectly well, which makes their alleged continued participation in scraping paywalled content a deeply cynical commercial calculation. They saw the cliff edge. The filings suggest they kept walking. The fear was that the greatest technological gold rush in a generation was being built on a foundation of sand, and that the foundation was being washed away by the very product it supported.
The fight goes to court
These admissions are not from a leak. They are from a court case. The explosive quotes are part of documents unsealed in the lawsuit brought by The New York Times against both Microsoft and its prominent partner, OpenAI. This is a legal war. It is a defining battle over intellectual property in the twenty first century, pitting some of the world’s oldest media organisations against its newest and most valuable technology firms. The central question appears straightforward. Can an AI company take the entire published archive of a news outlet, without asking for permission and without offering payment, and use it to build a hugely profitable commercial product. The answer is not simple.
The publishers' case is copyright. News organisations argue their work, from breaking news reports to long form investigations, is protected by copyright law which exists to ensure creators can control and be paid for what they make. They contend that scraping their websites on an industrial scale to train large language models is a direct, colossal and unlawful infringement of that fundamental right. The technology firms disagree. They claim 'fair use'. This is a legal concept which permits the limited use of copyrighted material without a licence, for purposes like commentary, criticism, research or scholarship. OpenAI and Microsoft argue that training their AI is a 'transformative' use, because the model does not reproduce the articles but instead learns from them to create something entirely new, and should therefore be sheltered by fair use principles.
This is where the unsealed filing becomes a powerful weapon. It is a gift. A huge one. The New York Times’s lawyers can now argue in court that Microsoft’s own senior staff privately agreed with the publisher’s central claim all along. The defence of 'fair use' becomes substantially harder to maintain when your own internal communications describe the activity as 'theft' and the 'largest theft of labor in human history'. It suggests the company knew its actions were ethically and legally dubious, yet proceeded for commercial advantage. The word 'theft' is poison for a fair use argument. Fair use rests on a degree of good faith, on creating something new and not harming the market for the original work. An internal admission of theft demolishes that posture, replacing the complex legal argument with a much blunter reality. They knew. They did it anyway.
The case is far from over. The ground, however, has shifted. This revelation dramatically strengthens the position of every publisher, artist, and author currently suing or thinking about suing AI companies for using their work without payment. The pressure is on. It forces Microsoft and OpenAI to consider settling the case, which would involve paying vast sums for content licences and setting a precedent for the entire industry. A courtroom defeat for the AI giants could be catastrophic for them, potentially forcing them to pay billions of pounds in damages and fundamentally altering how all future models must be built. The stakes are immense. The court must now weigh the complex 'transformative use' argument against a simple, damning admission from inside the machine itself.
So who pays for the future?
This revelation forces a choice. A painful one. Microsoft and OpenAI now face a clear fork in the road, with each path leading to a radically different future for artificial intelligence and the media. One route leads to negotiation tables and chequebooks. This is the world of expensive licensing deals. It would mean paying news organisations, perhaps billions of pounds collectively, for the right to use their archives as training material for large language models. The cost would be enormous. It would slow the AI gold rush, adding a huge new expense to a business model that has so far treated human knowledge as a free resource to be harvested. A future built on licensing changes the economics of AI, forcing developers to be more selective about data and creating a market where only the richest corporations can afford to build the most powerful models. The cost would eventually be passed on. Everyone would pay more.
The other path is to fight on in court. It is a huge gamble. If the courts ultimately reject the 'theft' narrative and side with the AI firms' fair use argument, the immediate consequences for publishers would be severe. Devastating, even. This is the 'doom loop' scenario that Microsoft's own executives feared. Without revenue from licensing, and with AI chatbots providing instant summaries of their work for free, subscription and advertising models would crumble. Local papers would go first. Then national ones would follow. The information ecosystem would shrink and degrade, leaving behind a wasteland of state propaganda, marketing copy and low quality user generated content. The AI models themselves would then begin to starve. They would be forced to train on their own synthetic output, a process researchers call model collapse, which leads to a gradual, spiralling decay in quality and coherence. They eat themselves. In this future, nobody pays for journalism, and everyone gets a worse internet populated by increasingly nonsensical machines.
This legal battle is now a defining moment for the entire AI industry. The gold rush mentality, the drive to build bigger and faster at any cost, has met its first real obstacle. For months the strategy has been to scrape first and ask for legal permission later. That strategy looks broken. It looks reckless. The 'theft' admission gives regulators and politicians a powerful reason to intervene, framing the debate not as an abstract copyright dispute but as a simple case of big technology taking things without paying. The government has so far preferred a light touch. This could change everything. An industry that privately calls its own core practices 'theft' is a difficult one for any politician to defend publicly. Officials in Whitehall and Brussels will be watching the New York Times case with intense interest. A court ruling that forces licensing deals would solve a political problem, creating a market based solution. A ruling against the publishers would almost certainly trigger calls for direct legislative intervention to protect the press. Intervention looks likely. The only question is what kind it will be.
Sources. Ars Technica: Microsoft exec called AI scraping the “largest theft of labor in human history”. TechCrunch: Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal.
Analysis. Drafted with AI assistance from the sources listed above and reviewed by an editor before publication. Jnews links to the organisations it writes about.

