'Astonishing Theft of Unprecedented Scale': OpenAI and Microsoft On The Record
news — 09.18.26
Court documents unsealed Thursday reveal Microsoft and OpenAI execs candidly admitting that their LLMs are designed to extract value from unpaid labor and circumvent paywalls. “Perhaps the largest theft of labor in human history," Microsoft’s own Director of Applied Science Dr. Brent Hecht called it. Hecht is an academic that has been in his role at Microsoft since 2023. His own prior research focuses on building “sustainable AI-content ecosystems,” but even as a Microsoft employee he described its AI strategy as a “doom loop" that could damage the economic foundations of the internet.
It’s an acknowledgment of the problem LLM PR copy usually dismisses. Externally they defend their unpaid use of copyrighted content, but internally, the damage and potential scope of appropriating this labor is a known and openly discussed issue. But they do it anyway.
In their own words
OpenAI quotes published in the brief reveal that the "labor theft" is intentional. Nick Ryder, VP of Research at OpenAI informed OpenAI co-founder Greg Brockman of a hack to "get around [the New York Times] paywall," to which Brockman responded "ah nice.”
That’s one of the core issues in the NY Times v. OpenAI lawsuit, the intentional subversion of the paywall to get access to articles to feed ChatGPT. That’s more than just scraping publicly available information, that’s breaking the lock to snatch content OpenAI didn’t pay for. Which follows the quote from an unnamed source at Microsoft that "almost no one intended for content to be used in this fashion, nor are they compensated for its use."
Smash and grab
That’s the broader issue outside this lawsuit: LLMs don’t exist without information, data, and content their creators didn’t pay for. The common counterargument is okay sure, but the news and article summaries from LLMs do include links back to their sources. That’s not a solution, and they know it.
From the brief, an unnamed OpenAI software engineer stated "no matter how prominently we show the links, users won’t click." And Nick Turley, OpenAI’s ChatGPT head¹ admits users have “no good reason to click” a source link given in a news or article summary. And they’re right, and again, this isn’t just a symptom they’re naming, it is, in their own words, part of the business model.
Killing culture
Strip mining the media they rely on, extracting value till there’s nothing left is all part of the game. OpenAI’s Turley continued, stating models like ChatGPT are "largely substitutive" and "will get more and more substitutive as they get better." Microsoft CEO Satya Nadella is later quoted, under oath, "[Chatbots] give you the information right there on the website on the AI platform versus needing to go to the underlying source."
These are the same executives arguing publicly and in court that their products are non-substitutive and fair use — all while privately putting the lie to their own words. OpenAI’s Jack Clark² said it plainly: the goal of ChatGPT is to ‘lead to us creating systems that substitute for the labor of the people that define the culture of society.’
Further Reading:
404 Media: ‘Doom Loop’: OpenAI and Microsoft Admits LLMs Are Destroying the Web and Built on Theft
Ars Technica: Microsoft exec called AI scraping the “largest theft of labor in human history”
¹: per his Linkedin page.
²: Jack Clark was OpenAI’s policy director from 2018-2020 when he left to co-found Anthropic.