SecurityTechInsider AI security & governance
EN/ NL
Governance

Court forced to destroy AI models? What to record now about your model provenance

US newspapers demand destruction of AI models trained on their journalism. What that demand means for the provenance and governance of your AI stack.

8 September 2026 4 min
Illustration for this article: Court forced to destroy AI models? What to record now about your model provenance. Optical fibre ends clustered together, each carrying a pinpoint of light in near darkness.
Organisations must now track AI model provenance and maintain the ability to retrain or switch providers if courts order destruction of infringing training data. Image: SecurityTechInsider — original editorial illustration

You must now document, per workflow, which models you deploy and what you know about their training data provenance, so you can switch providers or retrain rapidly if a court orders destruction of infringing material.

The prompt is an analysis of 8 September 2026 of model destruction as a remedy in copyright litigation, which argues that destruction orders are no longer theoretical but now appear as concrete demands in active US lawsuits. The Seattle Times and Newsday have sued OpenAI and Microsoft in federal court, explicitly requesting that the court order destruction of training datasets and models containing their journalism. In our assessment, the shift that matters operationally is this: "destroy the model" has moved from legal commentary into active procedure, and organisations that depend on trained models now face a concrete governance duty to track provenance and maintain the ability to switch or retrain.

What legal grounds are the newspapers using?

The newspapers base their destruction demand on copyright infringement. They allege that OpenAI and Microsoft collected their journalism—including paid articles behind paywalls—without permission and used it to train models such as GPT and Copilot. The New York Times made a similar demand in December 2023, citing 17 U.S.C. § 503(b), which allows courts to order destruction of infringing copies and tools. A separate case involving 35 regional and local publishers, collectively representing roughly 400 newspapers, characterises the practice as "industrial-scale" automated scraping combined with removal of copyright information and reproduction of articles by the models themselves.

How is the federal government responding?

The US Department of Justice has signalled support for the model developers, arguing that training on copyright-protected text constitutes fair use rather than infringement. This directly contradicts the newspapers' premise that training itself amounts to a violation. The outcome of this clash remains undecided. No court has yet ruled on whether a trained model is legally an infringing copy, a derivative work or a lawful tool; whether destruction of a deployed system is even operationally feasible; or how courts would weigh destruction against alternatives such as damages, licences or retraining on different data.

What operational risks does a destruction order create?

If a court orders destruction of training datasets or models containing specific publishers' work, the impact could extend across multiple layers of your AI stack:

  • Model availability — a provider may be forced to withdraw or retrain models, leaving you without the tool you depend on.
  • Workflow interruption — retraining or switching to an alternative model takes time and may require changes to your processes.
  • Data governance exposure — if your workflows use models trained on material later found to infringe, you may inherit liability or operational disruption.
  • Contract enforceability — existing agreements with providers may not protect you if they face a destruction order.
  • Audit and evidence burden — you may need to demonstrate rapidly what data your models were trained on and how you verified provenance.

Which concrete controls must you be able to demonstrate?

Regardless of how courts ultimately rule, these cases make training data provenance a practical governance duty. You must establish controls that let you answer these questions per workflow:

  1. Document the model and its purpose — record which model each workflow uses, the lawful basis for any personal data it touches, and the date of deployment.
  2. Establish audit rights with your provider — ensure your contracts give you access to evidence about training data sources, versions and any changes to model composition.
  3. Track model versions and lineage — maintain records showing which version of a model you deployed, when, and what training data it contained, since model names can mask version changes.
  4. Plan for rapid switching — assess whether you can move to alternative models or retrain on different data if your current provider faces a ban or destruction order.
  5. Verify data provenance claims — do not assume a provider's assurances about training data; establish what verification steps you can perform or require as a condition of use.

What tooling can support this, and what remains your responsibility?

Verification tools can make visible which models are deployed per task and what data provenance claims accompany them. Such tools cannot guarantee correctness or remove your obligation to make final judgements about acceptable risk. They support your governance process; they do not replace it. Your professional assessment of whether a model's training data provenance is acceptable, whether the legal risk is tolerable, and whether you can absorb the operational cost of a forced switch remains yours alone.

Sources: This article draws on reporting and guidance from Reuters, United States District Court (via CourtListener), MediaNama and Wikipedia.

Marit Halversen

Written by

Marit Halversen

Covers AI governance and regulatory design, with a focus on how compliance obligations land on architecture rather than on paperwork.