top of page

US Government Backs OpenAI in New York Times Copyright Case: Fair Use and AI Training

  • 3 hours ago
  • 5 min read

The U.S. Department of Justice has formally entered the copyright fight between OpenAI and The New York Times with a position that could influence the economics of generative AI training far beyond this single case.


In a brief filed in Manhattan federal court on September 2, 2026, the government backed OpenAI's argument that using copyrighted works to train large language models can qualify as fair use, describing generative AI training as highly transformative and linking the issue to scientific progress, national security, and U.S. economic competitiveness.


The filing is significant because it places the federal government's institutional weight behind a core legal theory used by OpenAI and other model developers, but it does not decide the lawsuit, bind the judge, or establish that every training practice involving copyrighted material is lawful.


THE DOJ FILING GIVES OPENAI A POWERFUL NON-BINDING LEGAL ALLY.


The immediate change is legal positioning: the government now supports the argument that model training can be transformative while leaving the court responsible for the actual fair-use determination.


The New York Times sued OpenAI and Microsoft in 2023, alleging that millions of copyrighted articles were used without authorization to train the models behind ChatGPT and that some outputs can reproduce or closely imitate protected material.


OpenAI has consistently argued that training is a separate technological use in which models learn statistical relationships, language patterns, and factual associations rather than operating as searchable archives of source works. The Justice Department's intervention reinforces that distinction at the policy level.


The legal effect remains limited. An amicus-style government filing can shape the court's analysis and signal the federal government's preferred interpretation, but the judge still has to apply copyright law to the record in the case.


........


Element

What the government position supports

What it does not establish

Training purpose

Training can be treated as highly transformative

Every dataset or training method is automatically lawful

Fair use

Copyright law can permit use of protected works for model development

A categorical exemption for AI companies

Policy interest

AI development is tied to innovation, security, and economic competitiveness

Policy considerations replace the statutory fair-use analysis

Effect on the case

OpenAI gains a significant institutional ally

The DOJ decides the outcome of the lawsuit


........


The distinction is important because the government's position is broader than a defense of one company, yet narrower than a rule that would immunize the AI industry from copyright liability.


··········


TRAINING, MEMORIZATION, AND OUTPUT INFRINGEMENT REMAIN SEPARATE LEGAL QUESTIONS.


A court can view the training process as transformative and still examine how data was acquired, whether protected expression was retained, and whether particular outputs substitute for the original work.


U.S. fair-use analysis is fact-specific and traditionally evaluates the purpose and character of the use, the nature of the copyrighted work, the amount used, and the effect on the potential market for the original.


For large language models, the most consequential issue is whether ingesting protected works serves a sufficiently different technical purpose from publishing or licensing those works to readers. OpenAI argues that the training process is transformative because the model uses source material to learn generalized relationships rather than to distribute the source material itself.


Publishers can still contest that characterization by focusing on market substitution, unauthorized copying during the training pipeline, memorization, or outputs that reproduce protected expression. Those claims are analytically different from a blanket assertion that any exposure of an AI model to copyrighted material is infringement.


This separation also matters operationally. A model developer may have a strong fair-use argument for training and still need controls for output similarity, dataset provenance, opt-outs, licensing obligations, and source-specific restrictions.


OpenAI's own public position has reflected this layered approach: it has defended training as fair use while also maintaining publisher opt-out mechanisms and working to reduce verbatim regurgitation.


··········


THE FILING COULD REDUCE LEGAL UNCERTAINTY WITHOUT ELIMINATING COPYRIGHT COSTS.


If courts increasingly accept transformative-training arguments, the economic pressure may shift from blanket permission for model training toward provenance, licensing strategy, output controls, and claims involving direct market substitution.


The government's intervention arrives as AI developers face multiple copyright lawsuits and as policymakers debate whether model training should require broad licensing regimes. A judicial approach favorable to transformative training would materially affect the cost structure of frontier-model development because compulsory licensing across web-scale corpora could be extremely expensive and operationally complex.


At the same time, a favorable training ruling would not make licensed data economically irrelevant. High-quality archives, structured professional data, real-time feeds, contractual access, and datasets with clear provenance can still have substantial value for performance, compliance, and risk management.


The same day, U.S. officials also urged G20 countries to preserve room for AI training on copyrighted works while protecting creators, reinforcing the broader policy direction behind the Justice Department's filing.


........


Stakeholder

Likely implication if the fair-use argument gains ground

OpenAI and Microsoft

Stronger defense against claims that training itself necessarily requires permission

Other AI developers

Lower risk of a universal licensing obligation, but continued exposure around provenance and outputs

Publishers

Greater incentive to focus claims on substitution, reproduction, licensing markets, and access restrictions

Creators

Rights disputes may concentrate more heavily on identifiable use, market harm, and reproducible expression

Enterprise AI buyers

Vendor copyright controls and indemnity terms remain relevant even if training receives broader protection


........


For enterprise customers, this means that a favorable legal trend for model developers should not be interpreted as permission to ignore copyright risk inside retrieval systems, custom fine-tuning pipelines, internal knowledge bases, or generated deliverables.


··········


THE NEW YORK TIMES CASE CAN SHAPE AI TRAINING ECONOMICS BEFORE IT PRODUCES A FINAL RULE.


The practical decision rule is to separate the legality of general model training from the narrower risks created by data acquisition, memorization, output reproduction, and market substitution.


The Justice Department's September 2 filing strengthens OpenAI's position because it frames generative AI training as a transformative use with strategic national value, and it gives courts a clear government view on a legal question that has remained unsettled across the industry.


Its importance should still be measured carefully. The filing is advocacy, not precedent. The court can reject parts of the government's reasoning, distinguish the facts of this case, or find that some conduct surrounding training and output generation deserves different treatment.


For model developers, the rational response is therefore continued investment in dataset provenance, output safeguards, contractual data access, and targeted licensing where the commercial or legal case supports it. For publishers and creators, the strongest claims may increasingly depend on demonstrable substitution, reproducible expression, or specific market harm rather than the mere fact that a protected work entered a training corpus.


If that legal direction holds, the economics of frontier AI will favor systems that can prove where data came from, control what models reproduce, and distinguish transformative learning from distribution of protected content.


FOLLOW US FOR MORE.


·····


DATA STUDIOS


·····


[datastudios.org]

Recent Posts

See All
bottom of page