Yovao News · The World, In Focus. From Local to Global, Never Miss a Beat

Unsealed Filings Reveal Microsoft Exec Called AI Scraping ‘Largest Theft of Labor in Human History’

Unsealed Filings Reveal Microsoft Exec Called AI Scraping ‘Largest Theft of Labor in Human History’

Unsealed court documents in a copyright lawsuit filed three years ago by The New York Times against OpenAI and Microsoft have exposed internal communications that characterize AI training practices as theft and an existential threat to journalism. The filings reveal that a senior Microsoft executive privately labeled the companies’ scraping activities as “the largest theft of labor in human history,” while OpenAI leadership acknowledged that their models posed a severe risk to the publications whose work was used to train them.

The newly released material details allegations that both firms bypassed paywalls without detection, constructed training datasets through mass scraping, and intentionally removed copyright notices from the data before it entered their models. Much of the unredacted content originates from The Times’ own legal briefs, as the underlying exhibits remain sealed, and the quotes are presented without their full original context.

These admissions challenge OpenAI’s fair use defense, which relies on the argument that using copyrighted material for training does not harm the market for the original work. However, internal Microsoft data indicates that its Copilot “answer engine” caused click-through rates to the New York Times domain to plummet by as much as 93% compared to traditional Bing search results. Brent Hecht, Microsoft’s Director of Applied Science, described this dynamic in a January 2024 presentation as a “doom loop” that would simultaneously degrade model performance and harm the broader web.

“It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain,'” a Microsoft document stated, as quoted in the filing. CEO Satya Nadella also testified earlier this year that any paywalled content used for training should be licensed. He added that if he had known OpenAI had scraped and trained on paywalled information, he would have required the company to retrain its models.

Other internal communications further undermined the fair use argument. Nick Turley, OpenAI’s Head of ChatGPT, wrote in internal messages that publishers face an “existential threat” from products that are “largely substitutive” and will become more so as they improve. OpenAI President Greg Brockman described the models as “excellent at news,” and Nadella conceded under oath that chatting with AI has substituted for users visiting the original websites to get information directly.

A Microsoft document warned of a “real risk” that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.” The scale of the copying is substantial: the documents show that OpenAI’s mid-training datasets contained over 91,692 copies of works from The New York Times, the Daily News, and the Center for Investigative Reporting. Additionally, a dataset derived from Common Crawl included more than 2 million documents from nytimes.com alone.

In a January 2023 internal memo, Hecht referred to the situation as “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history.” The filings also outlined how OpenAI and Microsoft exchanged training data through initiatives known as Project Taxi and Project Mango, with the latter allegedly containing copies of at least 160,903 unique works from news publishers.

Further allegations suggest OpenAI employees developed strategies to circumvent paywalls. When researcher Nick Ryder shared a method to bypass the New York Times paywall with Brockman, the president reportedly replied, “ah nice.” The companies also allegedly constructed datasets like WebText and WebText2 with a heavy reliance on scraped news content and pulled millions of articles from Common Crawl. Efforts were made to strip copyright notices from the training data because researchers did not want the models to output them to users. Representatives for Microsoft and OpenAI did not respond to requests for comment.

5 responses to “Unsealed Filings Reveal Microsoft Exec Called AI Scraping ‘Largest Theft of Labor in Human History’”

  1. Does anyone else find it ironic that these tech giants claim AI is safe while destroying the very people who fuel it?

  2. I knew scraping was aggressive, but removing copyright notices intentionally? That crosses from gray area into clear malice.

  3. Calling it ‘theft’ internally but suing for fair use externally? That is a blatant contradiction. Something smells fishy here.

Leave a Reply

Your email address will not be published. Required fields are marked *