Skip to content
goppo

News · AI summarised to understand what matters

← Back to news

Regulation & Society

Published on

Microsoft internally called AI training "the largest theft in history"

Unredacted filings in the New York Times' lawsuit against OpenAI and Microsoft reveal that executives at both companies privately described their AI training practices as "theft," while OpenAI's leadership acknowledged an "existential threat" to publishers.

  • openai
  • microsoft
  • direitos autor
  • chatgpt
  • nyt

Summary

Unredacted material made public in the lawsuit The New York Times has brought against OpenAI and Microsoft since December 2023 reveals that executives at both companies privately described their AI training practices as "theft," and that OpenAI's own leadership acknowledged its models pose an "existential threat" to the publishers and journalists whose work trained them.

In practice

According to the filings, Brent Hecht, Microsoft's director of Applied Science, wrote in a January 2023 internal memo that the content-gathering amounted to "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history." Nick Turley, OpenAI's head of ChatGPT, wrote internally that products like the chatbot pose an "existential threat" to publishers, being "largely substitutive" and likely to become "more and more substitutive" as they improve. OpenAI President Greg Brockman described the models as "excellent at news."

Microsoft's own internal data, cited in the filing, shows its Copilot "answer engine" caused click-through rates to The New York Times' domain to drop by as much as 93% compared to traditional Bing search. An internal January 2024 presentation, also by Hecht, describes the decline as a "doom loop" that would "hurt the performance of our models and the entire web at the same time."

Microsoft CEO Satya Nadella testified in a deposition this year that "anything that is paywalled should be licensed by anyone who wants to use it… for grounding or training," and said that had he known OpenAI scraped and trained on paywalled content, he would have "invoked [Microsoft's right to] require OpenAI to retrain its models."

The filings also detail how the companies allegedly bypassed paywalls undetected: when OpenAI researcher Nick Ryder told Brockman about a "hack to get around nytimes paywall," Brockman replied "ah nice." The documents also point to deliberate efforts to strip copyright notices from training data, since researchers "wouldn't want [the] model outputting" those notices.

The scale of the alleged copying is significant: OpenAI's mid-training datasets alone contain more than 91,692 copies of works published by The New York Times, the Daily News, and the Center for Investigative Reporting. A Common Crawl-derived dataset included more than 2 million documents from nytimes.com alone. Per the filing, OpenAI delivered its entire GPT-3 training dataset to Microsoft, and Microsoft supplied training data to OpenAI through two internal initiatives called Project Taxi and Project Mango — the latter assembled into a dataset containing copies of at least 160,903 unique works from news publishers.

What we still don't know

Much of this information comes from The Times' own brief, not the underlying exhibits, which remain sealed; the quotes are presented without their original context. OpenAI and Microsoft did not return requests for comment. The case continues against a backdrop in which courts have largely favored AI companies' "fair use" arguments, and in which the Trump administration filed a brief this month backing OpenAI's position.

Why it matters

  • The internal quotes directly cut against OpenAI's fair-use defense, which depends on the use not substituting for or harming the market for the original work
  • The documented scale — more than 91,000 copied articles and millions of documents from a single site — gives concrete numbers to a dispute that had mostly stayed abstract
  • The paywall-bypassing and copyright-stripping allegations open a distinct line of liability beyond the fair-use debate itself
  • The case could shape how courts and regulators treat similar lawsuits against other AI companies