
Short Answer: What did the US government just say about AI training and copyright?
The Trump administration filed a 20-page brief in The New York Times’ copyright lawsuit against OpenAI, defending the company’s unlicensed use of copyrighted material to train its AI models. The brief argues that restricting AI training under a “misunderstanding” of fair use would hurt American AI leadership and economic growth. It is not a ruling — the court is not required to follow it — but it signals where federal policy is leaning.
Quick Summary
- The brief was filed in The New York Times’ lawsuit against OpenAI, in the U.S. District Court for the Southern District of New York.
- It leans on Trump’s executive order about retaining U.S. “global leadership in artificial intelligence.”
- The core legal question is fair use: whether training an LLM on copyrighted work without permission is “transformative” enough to be legal.
- Courts have mostly sided with AI companies so far — Anthropic’s $1.5B settlement was for pirating books, not for training on them.
- The brief carries no legal force on its own, but it signals the administration’s policy leaning in a case still being decided.
TechCrunch first reported the filing on Tuesday.
What exactly did the government file, and why?
The Trump administration submitted a 20-page brief in The New York Times’ copyright lawsuit against OpenAI, arguing in favor of the company’s right to train ChatGPT on copyrighted material without a license. The brief does not come from a party to the case — it is an outside filing meant to influence how the court thinks about the issue.
Its central argument is about national competitiveness, not copyright doctrine specifically:
“The United States has a strong interest in continuing to develop a robust and competitive artificial intelligence industry that sets the standard for the practice and procedure of AI use globally… it is critical for the United States to ‘retain global leadership in artificial intelligence.’”
That language references an executive order President Trump signed last year. The brief goes further, warning that “constraining LLM development under a misunderstanding of fair use doctrine would thwart such creative and scientific progress while hindering American prosperity and economic mobility.”
What is The New York Times actually arguing?
LLMs like ChatGPT, Claude, and Gemini are trained on enormous datasets that include copyrighted books, articles, and other published work, typically pulled in without permission from the rights holders. The Times, along with a number of other publishers, argues that this practice is illegal when done without a license.
The legal fight centers on fair use — the copyright-law exception that permits using someone else’s work without permission in certain circumstances. The specific question here is whether training an AI model on copyrighted text is “transformative” enough to qualify, or whether it is closer to simply copying the work for commercial gain.
How have courts ruled on this so far?
Rulings to date have mostly favored AI companies, though the details matter. Last year, Judge William Alsup ordered Anthropic to pay a $1.5 billion copyright settlement to a group of writers — but notably, Anthropic was not penalized for training its models on their books. It was fined for how it obtained those books: through illegal shadow libraries used to pirate them.
Judge Alsup drew a distinction that’s likely to keep showing up in future rulings, comparing an LLM’s training process to a human reader:
“Like any reader aspiring to be a writer, Anthropic’s LLMs trained upon works not to race ahead and replicate or supplant them — but to turn a hard corner and create something different.”
In other words: training itself has generally been treated as defensible under fair use, while the sourcing of the training material has been where AI companies have actually gotten into legal trouble.
Does this brief actually decide anything?
No. The case is being tried in the U.S. District Court for the Southern District of New York, and the authors of the brief have no jurisdiction over the outcome — a judge still has to rule. But an intervention at this level from the federal government is a meaningful signal about where policy is heading, and it could shape how the court, and future cases, weigh the competing interests at stake.
Why does this matter for AEO and publishers?
If courts and policy continue trending toward “training on copyrighted work is fair use, full stop,” the practical reality for publishers doesn’t change much: your content is very likely already inside these models, and opting out of that going forward is not straightforward. Blocking crawlers can limit future training access on a going-forward basis — see our breakdown of which AI crawlers actually respect a disallow rule — but it does nothing to un-train a model on what it already ingested, and it does nothing to stop citation and retrieval-based tools like search-grounded answers, which typically aren’t governed by the same training-data legal questions at all.
That reframes where the real leverage is for a publisher. If you can’t meaningfully control whether your content is used to train a model, the thing you actually can influence is how that content gets cited when someone asks an AI system a question — which is a retrieval and attribution problem, not a training-data problem. It’s the same distinction we cover in why third-party citations matter more than your own content for AEO: getting named and linked when it counts is a different fight than trying to keep your work out of a training set in the first place.
It also raises the stakes on getting your own AI-facing infrastructure right rather than leaving it to chance. Tools like a well-maintained llms.txt file won’t stop training, and there’s no evidence yet that they influence citation either — but they’re a low-cost way to at least tell compatible systems what you consider your most important, most accurate content, in a landscape where the training question increasingly looks settled in AI companies’ favor.
Frequently Asked Questions
Did the US government rule that AI training on copyrighted work is legal?
No. The Trump administration filed a brief supporting OpenAI’s position in an ongoing lawsuit, but a brief is not a ruling. The presiding judge in the Southern District of New York still has to decide the case.
What is the lawsuit about?
The New York Times sued OpenAI, arguing it is illegal to train ChatGPT on the Times’ copyrighted articles without permission or a license. OpenAI argues the practice is protected under fair use.
Has any AI company actually been penalized for training on copyrighted work?
Anthropic paid a $1.5 billion settlement, but the penalty was for using illegal shadow libraries to pirate the books it trained on, not for the act of training itself. That distinction has been central to how courts have approached these cases.
Can publishers stop AI companies from training on their content?
Only partially, and only going forward. Blocking known AI crawlers via robots.txt can reduce future training access, but it cannot remove content already absorbed into an existing model, and it has no bearing on retrieval-based citation in AI search results.
What should publishers focus on instead?
Since training access is increasingly difficult to control or reverse, the more actionable lever is optimizing for citation and attribution when AI systems answer questions — the retrieval side of the equation, not the training side.
Source: TechCrunch.
About the author
Kai Williams
Kai Williams has been in marketing for years, with a long background in SEO before AEO had a name. He stepped into Answer Engine Optimization the moment AI started reshaping how people search, and has been tracking the shift ever since. At Prompt Insider, he covers AEO, AI marketing, and the future of search, breaking down what is changing and what brands need to do about it.


