“We decline plaintiffs’ invitation to transform run-of-the-mill copyright-infringement claims into DMCA claims,” Judge Eric Miller wrote for a unanimous Ninth Circuit panel.
The opinion came on Wednesday in Doe v GitHub, number 24-7700. Anonymous programmers who publish open-source code had sued GitHub, Microsoft and OpenAI over Copilot and Codex.
They argued the tools reproduce their code without attribution. That, they said, breached section 1202(b) of the Digital Millennium Copyright Act.
The panel affirmed the dismissal of that claim. The ruling leaves their two breach-of-contract claims running.
Copilot is not a search engine, and that decided it
Section 1202(b) bars intentionally removing or altering copyright management information. That means the author’s name, the copyright notice and the licence terms. In open-source code it sits in a header at the top of the file.
Miller held that removal requires an existing work to remove it from. “One who creates a new work and fails to include CMI cannot be said to have ‘removed’ or ‘altered’ anything,” the opinion says.
The court’s comparison is the clearest part of the ruling. A search engine retrieves and displays stored information, meaning copies of things that already exist.
“If Copilot functioned like a search engine and produced outputs that were identical to plaintiffs’ code but did not contain CMI, then plaintiffs might have a stronger claim,” Miller wrote. “But that is not what plaintiffs have alleged.”
By the programmers’ own account Copilot predicts a likely completion from statistical patterns. The court took that as describing generation rather than retrieval.
The court admits the statute fits AI badly
The certified question was whether section 1202(b) carries an identicality requirement. The panel said the label is wrong.
“Identicality” is “something of a misnomer”, Miller wrote. The DMCA does not demand literal sameness, and he called the label a gloss on the statutory words rather than a separate element.
That matters for the next case. The defendants themselves conceded that two works need not be identical.
The court agreed. Minor cosmetic changes will not protect someone who substantially reproduces a work and strips the notice.
The opinion also concedes the awkwardness of its own authorities. Its examples come from print media, such as defacing the title page of a book.
Those “do not map neatly onto the emerging digital technologies like artificial intelligence to which the DMCA’s protections also apply”, Miller wrote.
The bigger claim was given away in open court
The programmers ran two theories. The panel decided only the output theory, that Copilot emits memorised training data without the accompanying notices.
The other was the input theory, that the defendants stripped the notices from the code before feeding it into training. That is the claim about training data itself, and the panel refused to hear it.
The record shows why. At the hearing on the first motion to dismiss, the district court asked a direct question: did copying training data into Copilot breach the licences’ attribution requirement?
Counsel for the programmers answered “Perhaps it doesn’t.” The judge then observed that the “complaint is not about training. It just isn’t.”
A later order stated that the “Plaintiffs do not allege they were injured by Defendants’ use of licensed code as training data”. Nobody corrected it, orally or in writing.
The Ninth Circuit held the theory forfeited. The panel never reached the question everyone else is litigating over AI training, because the programmers had not properly raised it.
The programmers won on standing
The defendants argued the programmers could not show their own code, rather than some other contributor’s, was at risk. The panel disagreed.
The complaint cites academic research finding that large language models will sometimes “emit the memorized training data verbatim”, a phenomenon that “will likely get worse as models continue[] to scale”.
It also points at GitHub’s own duplicate-detection feature, which lets users block suggestions matching public code at verbatim snippets of 150 characters or more. The court treated the existence of that filter as some evidence that Copilot does emit identical copies.
That was enough at the pleading stage. The panel was explicit that at summary judgment the programmers would have to produce actual evidence.
Why the DMCA route was worth the fight
The opinion sets out the arithmetic plainly. The DMCA permits up to $25,000 per violation. Ordinary copyright statutory damages stop at $30,000 per work.
A tool answering millions of prompts generates violations, not works. Miller wrote that letting the claim through would supplant ordinary copyright law and expose defendants to “potentially ruinous liability”.
Courthouse News reports the programmers sought more than $9bn. The panel expressly declined to say whether Copilot’s output infringes in the ordinary way.
Publishers lined up against the AI companies
The amicus filings show who thinks this question matters. The Authors Guild, the Association of American Publishers, the News/Media Alliance and the International Association of Scientific Technical and Medical Publishers all filed, as did the Authors Alliance and a group of intellectual property law professors.
On the other side sat the Electronic Frontier Foundation and Public Knowledge, plus the Chamber of Progress, the Computer and Communications Industry Association and ACT The App Association.
The EFF argued that a broad reading would let rights holders sue “artists making remixes based on older works, teachers adapting works for a classroom presentation, engineers reverse engineering code to understand it better, and search engines that help us all navigate the web”.
Joe Mullin published the group’s response under the headline “Victory”. Open-source developers and digital rights groups are usually on the same side, and here they were not.
What is still alive
Two breach-of-contract claims remain before Judge Jon Tigar in the Northern District of California. Those turn on the open-source licences themselves rather than on federal copyright law.
The open-source lawyer Heather Meeker notes that licence notice requirements generally attach to redistribution rather than to training. Licences that grant use for any purpose may offer less protection than developers assumed.
The same provision failed in July, when a court said others could scrape Google in a suit Google had brought. Section 1202 has been a weak instrument against AI systems.
The money elsewhere has been larger. Anthropic settled a book-piracy case for $1.5bn in July, a fight about acquiring the training data rather than about the output.
Results outside the United States have not followed the same line. A German court found the AI music tool Suno broke copyright, a first for Europe.
The Justice Department sided with OpenAI in the publishers’ copyright fight this month. Code carries its own difficulty, as the 28-year fight over who owns Linux showed.
How it got here
The Joseph Saveri Law Firm and the lawyer Matthew Butterick filed the class action in the Northern District of California in 2022, as docket 4:22-cv-06823. Two rounds of dismissals and amendments cut it to three claims.
Jesse Panuccio of Boies Schiller Flexner argued for the programmers. Lisa Blatt of Williams and Connolly and Christopher Cariello of Orrick argued for the companies.
The panel heard argument in San Francisco on 11 February and ruled seven months later. Miller sat with Senior Circuit Judge Sidney Thomas and District Judge Stanley Blumenfeld Jr, who was sitting by designation from the Central District of California.
Get the TNW newsletter
Get the most important tech news in your inbox each week.