
Newly unsealed court filings have pulled back the curtain on private conversations inside Microsoft and OpenAI that read less like corporate memos and more like confessions. According to documents made public this week and reported by The New York Times and TechCrunch, employees at both companies debated internally whether scraping news content to train artificial intelligence models amounted to outright theft. The disclosures have reignited the AI training data controversy that has simmered since The New York Times first sued the two companies in late 2023, and they now sit at the center of one of the most closely watched copyright cases in the industry.
Key takeaways
- Microsoft’s director of Applied Science, Brent Hecht, privately called AI scraping “the largest theft of labor in human history” in a 2023 internal document.
- Internal Microsoft filings describe a “doom loop” in which AI tools like Copilot cut traffic to publishers, undermining the very content used to train the models.
- OpenAI president Greg Brockman replied “ah nice” when told about a plan to bypass The New York Times paywall.
- The Times’ lawsuit, filed in late 2023, has since been joined by eleven other publishers, and OpenAI has had to preserve 20 million ChatGPT conversation logs for the case.
- Judge Sidney Stein of the Southern District of New York is now weighing summary-judgment motions as more sealed material comes to light.
Internal Concerns Over AI Data Scraping at Microsoft and OpenAI
Employees at both companies raised alarm bells about their own AI training practices long before the lawsuit became public knowledge. The unsealed filings show that staff weren’t simply following orders — they were questioning, in writing, whether what they were building was ethically sound.
Microsoft Employees’ Debate Over ‘Labor Theft’
In a January 2023 internal memo, Hecht wrote that “millions of people around the world will soon consider large models hoovering up all their work” to be “an astonishing theft of unprecedented proportions.” That framing, according to the unsealed briefs, ran alongside internal discussion at Microsoft about whether AI scraping represented the largest theft of labor in human history. It’s a striking admission for a company that would go on to defend its training practices in court as lawful and transformative.
Warnings of a ‘Doom Loop’ and Paywall Workarounds
Hecht’s warnings extended beyond that single memo. In a separate internal presentation dated January 2024, he wrote that “large AI models are a product that destroys its supply chain,” describing what internal filings call a “doom loop” — a cycle in which AI answer engines like Copilot reduce traffic to the news sites that supply their training data, degrading the quality of future models. Microsoft’s own data reportedly showed click-through rates to The New York Times’ domain falling as much as 93% compared with traditional Bing search results, a decline the company itself described as threatening “the economic foundations of its essential suppliers.”
Over at OpenAI, the paywall problem was apparently handled more directly. According to the filings, a staffer told OpenAI president Greg Brockman about building a “hack” to get around the NYT paywall. Brockman’s reply, preserved in the court record, was simply: “ah nice.” That exchange has become one of the most quoted lines in the unsealed material, largely because it undercuts the idea that the paywall workaround was accidental or peripheral.
The Legal Battle Over Copyright and AI Training
The OpenAI Microsoft lawsuit has grown considerably since it was first filed, expanding from a single plaintiff into a coalition of publishers demanding accountability over how their journalism ended up training some of the world’s most valuable AI models.
New York Times Lawsuit and the Publishers Who Joined
The New York Times filed its suit against Microsoft and OpenAI in late 2023, arguing the companies had violated copyright law by training generative AI systems on its journalism without permission or payment. Eleven other publishers have since joined the litigation, turning what began as a single news organization’s grievance into a broader industry challenge to how AI firms source their training material.
Judge Sidney Stein and the Order to Preserve ChatGPT Logs
The case now sits before Judge Sidney Stein of the Southern District of New York, who is weighing summary-judgment motions as more documents get unsealed. As part of the litigation, OpenAI has been ordered to preserve 20 million ChatGPT conversation logs — a requirement that underscores just how far-reaching the discovery process has become in this dispute over AI model copyright issues.
Corporate Positions and Technical Arguments
Both companies maintain that their AI training practices are legally defensible, even as their own internal communications complicate that argument.
Microsoft’s Response and Nadella’s Licensing Testimony
Microsoft has moved to distance itself from Hecht’s memos, saying they reflect one employee’s personal analysis rather than official company policy. The company said Hecht, who also held a post at Northwestern University, was not a decision-maker and was specifically employed to “present divergent and asymmetric perspectives.” That framing matters legally, since it separates the company’s public defense from the blunt language used internally.
Yet CEO Satya Nadella’s own deposition testimony cuts closer to the heart of the dispute. In his testimony, he stated that “anything that is paywalled should be licensed by anyone who wants to use it” when it comes to training or grounding AI systems, adding that if he had known OpenAI was using paywalled content for training, he would have invoked Microsoft’s right to force OpenAI to retrain its models. A company spokesman later said Nadella “spoke to broad principles” about how people find and consume information — a clarification that hasn’t stopped the testimony from becoming a central exhibit in the case. The comments speak directly to the debate over paywalled content AI training, an issue that now sits at the core of the publishers’ legal argument.
OpenAI’s Fair Use Defense and Internal Warnings
OpenAI’s legal position rests on fair use: the company argues its training process transforms articles into something new rather than substituting for the originals, a distinction that has generally favored AI firms in prior court rulings. However, comments from OpenAI’s own employees undermine that argument: Nick Turley, the former head of the ChatGPT team, wrote in June 2023 that AI represented an “existential threat” to publishers, and subsequently observed that AI products “will get more and more substitutive as they get better.” An OpenAI engineer separately observed in February 2023 that “no matter how prominently we show the links, users won’t click” — a finding that works against the claim that chatbots drive meaningful traffic back to news sites.
The most pointed internal warning came earlier, in 2020, when then-policy director Jack Clark wrote to Brockman and Sam Altman that OpenAI was “creating systems that substitute for the labor of the people that define the ‘culture’ of society,” and risked becoming “the symbol of how Silicon Valley is thoughtlessly stepping into other parts of life and leaving a mess on the carpet.” Clark later left OpenAI to co-found Anthropic, which referred a request for comment on the matter back to OpenAI.
Steven Lieberman, who represents the New York Daily News and seven other papers in the case, said the unsealed material shows “the world can see what OpenAI and Microsoft thought all along about the fairness of their own behavior.” Neither OpenAI nor Microsoft has publicly responded to that characterization, and The Times itself declined to comment on the filings to its own reporters.
Taken together, these disclosures suggest the courtroom fight is no longer just about whether AI training counts as fair use in the abstract — it’s about whether the companies’ own words, written years before the lawsuit reached this stage, will end up shaping how judges define fair use for the entire industry going forward.
FAQ
What labor-related concerns did Microsoft employees express regarding AI data scraping?
Microsoft employees discussed whether OpenAI’s use of news articles was “the largest theft of labor in human history” and warned of negative impacts on AI model quality, describing a “doom loop” that could degrade the very models built on scraped content.
Who filed the lawsuit against Microsoft and OpenAI and when?
The New York Times filed a lawsuit in late 2023, later joined by eleven other publishers.
What is Microsoft CEO Satya Nadella’s position on using paywalled content for AI training?
Satya Nadella testified that paywalled content should be licensed by anyone who wants to use it for training or grounding AI systems.
How does OpenAI defend its use of copyrighted material in AI training?
OpenAI argues its training constitutes fair use by transforming articles into new work rather than substituting for the originals.
Article produced with the assistance of artificial intelligence and reviewed by the editorial team.

4 hours ago
31








English (US) ·