A Microsoft scientist called AI training the largest theft of labor

Microsoft's Director of Applied Science called AI training perhaps the "largest theft of labor in human history", in documents unsealed on Thursday. Another Microsoft file calls the result a "doom loop". Its own data shows Copilot cut news click-throughs by up to 94%.


Aged printed and handwritten documents with official seals, spread across a bed and photographed in black and white.
Image Credits Credit: MicheleAroundTheWorld on Unsplash

A 92-page brief unsealed on Thursday opens by quoting the same Microsoft employee twice. Dr Brent Hecht, the company’s Director of Applied Science, wrote that people would come to see large models hoovering up their work as “an astonishing theft of unprecedented proportions”. In a second document he called it perhaps the “largest theft of labor in human history”.

Hecht is a research scientist. He is not a board member and he is not the chief executive, which matters as the phrase travels.

The News Plaintiffs’ combined summary judgment brief sits in the consolidated OpenAI copyright litigation before Judge Sidney H. Stein. Redactions covered most of its damaging passages until Thursday.

Jason Kint of the trade body Digital Content Next spotted the unredacted version. George Hammond and Stephen Morris reported it that day for the Financial Times, Scott Nover and Gerrit De Vynck for the Washington Post, Jason Koebler for 404 Media, and Karen Weise and Mike Isaac for The New York Times, which is also a plaintiff.

The quotations below come from the filing itself.

The plaintiffs are not only The New York Times. They include the Daily News group and the Center for Investigative Reporting. They also include Ziff Davis, whose titles run to CNET, ZDNET, PCMag and Mashable. The people suing are, in other words, this publication’s direct peers.

The doom loop, in Microsoft’s own numbers

The brief quotes a Microsoft document describing what the company had built. Its AI content strategy, the document says, “has started a ‘doom loop’” that will damage both its own models and the entire web. The document calls the position unusual. An end product now threatens the economic foundations of its own suppliers.

The figures behind it are Microsoft’s own. The brief cites company data on click-through rates falling 87% to 93% for the Times’s websites, 83% to 91% for the Daily News group’s, and 51% to 94% for Ziff Davis’s. The comparison is Copilot against traditional Bing search.

Hecht wrote that document soon after this lawsuit was filed. The systems work, he explained, by replicating the patterns in their training content. There is no way of passing economic value back down the supply chain. That, he wrote, “necessarily threatens the economic stability of those who create the content”. A ruling for his own employer, he added, would arguably “make a complete mockery of the idea of ‘fair use’”.

Other Microsoft material in the brief goes further. One document warns of a “real risk” that the technology could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained”. Another puts it in eight words: “LLMs are a product that destroys its supply chain.”

‘Users won’t click’

The sharpest line in the filing belongs to an OpenAI software engineer the brief does not name. “No matter how prominently we show the links, users won’t click.”

Nick Turley, who runs ChatGPT, wrote that publishers face an “existential threat”. The products, he said, are “largely substitutive, period” and will get more so as they improve. Jack Clark, OpenAI’s policy director, put it another way. The company’s work would increasingly mean building systems that substitute for the labour of the people who define a society’s culture.

Satya Nadella agreed under oath that chatbots had substituted for going to the underlying source. Internal OpenAI documents call ChatGPT “the modern newsstand”.

What they knew they were training on

OpenAI’s VP of Research gave the brief a sentence its lawyers will enjoy. “We train our networks to memorize the training data. That’s their objective.” By November 2019 the company was worrying internally about regenerating copyrighted works.

The brief puts numbers to the copying. OpenAI’s mid-training datasets hold more than 91,692 copies of the plaintiffs’ works. WebText2 held at least 6,552 from the Times, 18,609 from the Daily News group and 66,780 from Ziff Davis.

OpenAI separately obtained the New York Times Annotated Corpus, over 1.8 million articles from 1987 to 2007, under a licence limited to non-commercial research.

Around 2017, Greg Brockman wrote that he was “deeply motivated by the gazillions” he hoped to make from commercialising OpenAI’s technology.

Then there is the exchange the coverage keeps returning to. Nick Ryder told Brockman about “a hack to get around nytimes paywall”. Brockman replied: “ah nice.” OpenAI’s own store lists Custom GPTs named “Bypass Paywall” and “Remove Paywall”. Nadella testified that “anything that is paywalled should be licensed”.

Horse trading

The brief gives a whole section the title “Horse Trading”. Microsoft scraped for Bing, the filing says, then passed that content to OpenAI. OpenAI handed Microsoft the entire GPT-3 training dataset. Two joint initiatives, Taxi and Mango, moved data back, and the Mango dataset alone held at least 160,903 unique plaintiff works.

What the government told the same court

Sixteen days before the unsealing, the United States filed its own statement of interest in the same litigation. It argues that training on copyrighted text is fair use. TNW covered it at the time, and the 20-page filing reads differently now.

Its core argument is that training and output are separate uses. Copying a work to train a model reveals nothing to the public, so it cannot substitute for the original. Market harm counts only where the output is substantially similar to the source. General competition does not qualify.

It also frames licensing as an antitrust problem. Only the biggest technology companies could afford the fees, it argues, which would become “large subsidies for old mainstream media companies”.

What the brief does not prove

This is the plaintiffs’ document. Lawyers building a case chose every quotation in it, and the exhibits underneath are still sealed.

Microsoft has answered the central quotation. Hecht’s words “reflect one employee’s individual perspective, are not a legal analysis”, the company told the Financial Times. It said Nadella’s testimony spoke to broad principles rather than reaching a conclusion on the copyright questions. OpenAI did not respond to the FT. It has previously called the case an attempt at “an undeserved payday at the expense of progress that benefits everyone”.

The defendants have not conceded the legal point either. The brief quotes Microsoft’s own economist expert, Dr Tucker, drawing the distinction their lawyers will rely on. Search engines return a ranked list of links. Grounded models synthesise information from retrieved sources. Whether that synthesis counts as substitution in the sense copyright recognises is the question in front of the judge.

Courts have so far leaned towards the AI companies. Anthropic settled its book piracy case for $1.5bn rather than test it, and publisher suits against Google over Gemini are still running.

Europe is asking the opposite question

Brussels spent this month asking publishers whether Google’s AI opt-out works, on the assumption that a publisher should be able to refuse. Washington has told a court that letting them refuse would be an antitrust harm.

The doom loop has a payroll attached on both sides of the Atlantic. Reach, publisher of the Daily Mirror, cut 220 editorial jobs this month and pointed at AI summaries eating its traffic.

Judge Stein now has both documents. One says the training was transformative and the harm is not the kind copyright recognises. The other is the defendants, in their own words, calling it theft. The summary judgment ruling settles which reading the law accepts.

Get the TNW newsletter

Get the most important tech news in your inbox each week.

Also tagged with


Published
Back to top