Seattle Times and Newsday Sue OpenAI and Microsoft: Paywalled Journalism, AI Training Data, Fair Use, and Model Destruction Claims
- 21 hours ago
- 6 min read

The Seattle Times Company and Newsday LLC filed a 38-page federal complaint on September 4, 2026 against OpenAI and Microsoft, escalating the dispute over whether generative-AI companies can copy journalism for training, retrieval, and output generation without a license. The case is especially consequential because the publishers say the alleged copying reached behind subscription paywalls and because the remedies they request extend beyond damages to the possible impoundment or destruction of datasets and models incorporating their work.
The complaint is an allegation, not a court finding. It identifies ChatGPT, Microsoft Copilot, and Bing AI features as products allegedly built or operated using the publishers’ journalism, while OpenAI continues to rely on fair-use arguments for model training and Microsoft has said it was surprised by the suit and remains open to discussing solutions for local journalism.
The technical issue is therefore broader than whether an AI system can quote a paragraph. The litigation targets the full content pipeline: acquisition, paywall access, dataset construction, model training, retrieval, generated answers, attribution, and the economic substitution that publishers say occurs when an AI answer satisfies a user without a visit or subscription.
··········
THE COMPLAINT TARGETS THE FULL AI CONTENT PIPELINE.
The publishers frame the dispute as a chain of alleged copying events rather than a single training-stage question.
The suit was filed in the U.S. District Court for the Southern District of New York and is reported as case 1:26-cv-07644. The plaintiffs allege that OpenAI and Microsoft scraped hundreds of thousands of articles, including material accessible only to paying subscribers, and incorporated that journalism into systems used to train or operate generative-AI products.
The complaint also points to output behavior. It alleges that the systems can reproduce passages verbatim or closely paraphrase reporting, sometimes while omitting or altering titles, bylines, copyright notices, or other rights-management information. One cited example involves an 88-word passage from Seattle Times reporting on the Boeing 737 MAX, which the publishers present as evidence that protected expression can be recoverable from the systems under certain prompts.
........
Stage | Publishers’ allegation | Technical or legal issue |
|---|---|---|
Acquisition | Articles were scraped from publisher websites, including paywalled pages | Authorization, access controls, terms of service, and copying |
Dataset construction | Publisher content was incorporated into datasets used to train or operate AI products | Reproduction rights, dataset provenance, and licensing |
Model training | Copyrighted journalism contributed to model development | Whether training is transformative fair use or infringement |
Retrieval and generation | Systems can reproduce or closely paraphrase protected reporting | Memorization, retrieval, substantial similarity, and market substitution |
Attribution and metadata | Outputs may remove or alter titles, bylines, or copyright-management information | DMCA-style rights-management claims and source integrity |
Commercial effect | AI answers can reduce visits and subscription demand | Market harm, licensing value, and competition with original journalism |
........
··········
FAIR USE WILL TURN ON PURPOSE, TRANSFORMATION, AND MARKET HARM.
The central defense is expected to collide directly with the publishers’ argument that generative systems can substitute for the originals.
OpenAI has consistently argued that training on publicly available material is permitted by fair use and that models learn statistical relationships rather than functioning as databases of copied works. That theory focuses on the purpose and transformation of the training process, the amount of protected expression retained, and whether generated outputs are ordinarily substitutive.
The publishers’ complaint attacks that framing at several points. Paywalled material is economically different from freely accessible web content because access itself is sold. If protected articles were obtained despite subscription controls and then reproduced or closely paraphrased by AI products, the plaintiffs can argue that both the acquisition mechanism and the downstream market effect weigh against fair use.
The case also arrives while the U.S. government has taken a more AI-friendly position in separate litigation involving The New York Times, arguing that model training can qualify as fair use and warning that an overly restrictive interpretation could impair U.S. AI development. That position is not binding on the court here, but it makes the Seattle Times–Newsday case part of a broader effort to define where copyright ends and industrial-scale machine learning begins.
Microsoft’s role adds another layer. The complaint does not treat the dispute as only an OpenAI training question; it also targets Microsoft products and integrations, including Copilot and Bing AI features. Liability therefore may depend on which company copied which works, which datasets or model outputs were used, how product-level retrieval operates, and whether each defendant can show an independent legal basis for its conduct.
··········
THE MODEL-DESTRUCTION REQUEST MAKES THE REMEDY TECHNICALLY UNUSUAL.
The publishers are asking for relief that could reach datasets and deployed model artifacts, not only monetary compensation.
The complaint seeks damages, injunctive relief, and impoundment or destruction of copies, datasets, or large-language models that incorporate the plaintiffs’ copyrighted works or derivatives. No judge has ordered such destruction; it is a requested remedy, and whether it is legally available or technically proportional will be heavily contested.
For modern AI systems, destruction is not a simple delete operation. A training corpus may be sharded across storage systems, transformed into deduplicated or filtered datasets, used to generate checkpoints, distilled into later models, or combined with retrieval indexes and fine-tuning data. Removing one publisher’s material after training can therefore involve very different technical tasks depending on where the content entered the pipeline and whether the claimed protected expression remains recoverable.
........
Possible remedy | What it could require | Technical difficulty |
|---|---|---|
Delete source copies | Remove publisher files from retained raw or processed corpora | Relatively bounded if provenance is complete |
Destroy training datasets | Delete dataset versions containing the articles | Harder when datasets are deduplicated, transformed, or distributed |
Remove retrieval indexes | Rebuild search, vector, or RAG indexes without the protected material | Feasible but operationally expensive at scale |
Model unlearning | Reduce a model’s ability to reproduce specific protected content | Technically difficult to verify and may affect unrelated behavior |
Destroy model checkpoints | Retire models trained on the disputed material | Potentially extreme if the material is a tiny fraction of a very large corpus |
Future-use injunction | Block unlicensed ingestion or reproduction going forward | Requires durable provenance, access-control, and output safeguards |
........
The distinction between a dataset and a trained model will therefore matter. Copyright law can order impoundment or destruction in some infringement contexts, but applying that remedy to a model whose parameters were influenced by billions or trillions of training tokens raises questions of causation, proportionality, traceability, and technical feasibility that conventional media cases rarely confront.
··········
THE CASE COULD FORCE AI COMPANIES TO EXPOSE HOW PUBLISHER DATA MOVES THROUGH THEIR SYSTEMS.
The discovery process may become as important as the eventual fair-use ruling because the publishers need to connect specific copyrighted works to specific datasets, models, retrieval systems, and outputs.
If the litigation advances, the most consequential evidence could include dataset manifests, crawler behavior, paywall handling, filtering rules, model versions, retrieval logs, memorization tests, licensing decisions, and internal assessments of whether generated answers substitute for visits to publisher websites. Those records can determine whether the dispute is about incidental exposure to web-scale data or systematic ingestion of valuable subscription journalism.
The economic structure is also shifting. Publishers increasingly negotiate direct AI licensing agreements precisely because their archives contain high-quality, time-stamped, edited information that can improve training, retrieval, and answer freshness. A ruling that treats unlicensed ingestion of paywalled journalism as legally distinct from ordinary public-web training would increase the bargaining value of those archives and could push model developers toward more explicit provenance and licensing systems.
A broad victory for OpenAI and Microsoft, by contrast, could reinforce the argument that transformative training is generally protected even when the source material is commercially valuable, leaving output-level reproduction and access-control circumvention as the more important boundaries. A broad publisher victory could make provenance, exclusion controls, licensing status, and post-training unlearning much more central to production AI architecture.
The immediate case is therefore not a judgment that OpenAI or Microsoft infringed copyright, and the requested destruction of models is not an existing court order. Its significance is that two major regional publishers are asking a federal court to connect copyright remedies directly to the technical artifacts of generative AI. That forces the litigation to confront a question the industry has largely deferred: when protected journalism enters a model pipeline without a license, what exactly must an AI company be able to identify, remove, compensate for, or prove was transformed?
·····
FOLLOW US FOR MORE.
·····
·····
DATA STUDIOS
·····
[datastudios.org]




