BusinessSep 2, 2026, 6:24 PM

US DOJ Backs AI Training Under Fair Use Doctrine

TickrWire Editorial Desk·Sep 2, 2026, 6:24 PM·3 min read AI-assisted, human-reviewed

Reported by The Decoder: US Department of Justice backs fair use for AI training in landmark copyright case. Analysis and context written by TickrWire.

30-second summary

The United States Department of Justice filed a brief arguing that training artificial intelligence models on copyrighted text constitutes fair use, directly opposing a previous stance from the US Copyright Office.

TickrWire
US DOJ Backs AI Training Under Fair Use Doctrine
Key takeaways
  • The US Department of Justice filed a legal brief supporting fair use for AI training data.
  • The filing counters a previous report from the US Copyright Office regarding mass copyright infringement.
  • The legal argument distinguishes between the internal training phase of models and their external text outputs.
  • Political controversy surrounds the dismissal of the former Copyright Register following her agency report.
Full story

The United States Department of Justice has officially intervened in ongoing legal battles between publishers and technology firms, siding with artificial intelligence developers. In the consolidated class action lawsuit featuring The New York Times, the department filed arguments stating that the process of training language models on copyrighted text falls under the legal protection of fair use. This position introduces a major federal voice into the complex debate over how intellectual property laws apply to modern machine learning systems.

The litigation began when The New York Times filed a lawsuit in a Manhattan federal court against OpenAI and Microsoft. The complaint alleged that millions of copyrighted articles were ingested without authorization to train advanced systems like GPT-4, ultimately creating tools that rival the publisher in the information market. The news organization demanded massive financial damages and requested the destruction of any models that incorporated its reporting. This high-stakes legal confrontation is viewed by experts as a critical test for the future of digital copyright enforcement.

At the core of the government argument is the distinction between the training phase and the final output generated by the models. The department notes that while entire works are copied into memory during training, those source materials are never published or distributed directly to the public. Furthermore, the generated outputs typically lack substantial similarity to the original texts, meaning that broad claims of market harm fail to recognize the technical reality of how these models function. The filing emphasizes that treating training and output as a single prohibited act is legally flawed.

To illustrate this perspective, the legal filing references author Joan Didion studying Ernest Hemingway by manually copying his stories during her youth. The department suggests that under alternative legal theories, such human learning practices could face liability if the subsequent original writings are deemed derivative. Comparing this to technology, the brief argues that requiring licensing every time an algorithm internalizes stylistic patterns would fundamentally alter how new works are created, mirroring human cognitive inspiration.

Beyond traditional analogies, the filing emphasizes the broader societal and creative utility of artificial intelligence tools. The government asserts that people utilize these systems to generate original works, and imposing strict liability across the board would stifle innovation rather than protect it. Interestingly, the brief points out that even journalists at the plaintiff newspaper have utilized automated tools for drafting and editing tasks, highlighting the ubiquity of the technology across modern creative industries.

However, these arguments face significant pushback regarding the sheer scale of commercial data harvesting. Critics and regulatory bodies point out a vast difference between an individual human reading or copying texts for personal education and a multibillion-dollar corporation utilizing vast datasets to build commercial products. This exact tension was highlighted previously by the US Copyright Office, which issued a report concluding that automated systems process data at a speed and scale that exceeds traditional human creation, making blanket fair use inappropriate for competitive commercial applications.

The Department of Justice submission directly attacks the validity of that prior regulatory report, arguing that the assessments carried no binding legal authority and ignored established case-by-case statutory analysis. The timing of these regulatory clashes adds further political context, as the former head of the Copyright Office was dismissed by the administration shortly after publishing the agency report. As the federal government itself remains divided on the matter, courts will ultimately decide how copyright statutes apply to the infrastructure powering modern machine learning.

Why this matters
Developers

The DOJ stance provides strong legal backing for continued training practices without mandatory blanket licensing.

Businesses

Commercial AI providers gain a powerful federal ally in ongoing high-stakes copyright litigation.

Investors

Legal risk profiles for foundational model developers could shift significantly depending on how federal courts rule.

Everyone

The outcome of this dispute will determine how artificial intelligence interacts with human journalism and art.

Glossary
fair use
A legal doctrine that permits limited use of copyrighted material without acquiring permission from rights holders.
Sources · 1
Read next
More stories
5 Free Courses to Go From LLM Beginner to PractitionerAI Tools

5 Free Courses to Go From LLM Beginner to Practitioner

A structured free course pipeline takes learners from neural network fundamentals to deploying production-grade LLM applications, curated by an AI educator.

Free Transcription with SpeakrOpen Source

Free Transcription with Speakr

Speakr is a free, open-source, and self-hosted alternative to commercial transcription services like Otter.ai, offering local data privacy, swappable AI backends, and automated workflows.

Path to Astra: critical capabilities and frontier safeguardsSecurity

Path to Astra: critical capabilities and frontier safeguards

OpenAI’s Astra model is the first to meet the Critical cybersecurity capability threshold under its Preparedness Framework, enabling autonomous discovery of unknown vulnerabilities and exploit chains.

Healthcare organizations can now connect EHR and additional industry data to ChatGPTAI Tools

Healthcare organizations can now connect EHR and additional industry data to ChatGPT

OpenAI has added an Epic EHR integration and a Healthcare Public Data plugin to ChatGPT for Healthcare, letting clinicians access patient records and official datasets in a single workspace.

Polimill builds Japan's next-generation public AI infrastructureAI Tools

Polimill builds Japan's next-generation public AI infrastructure

Polimill introduced QommonsAI, an OpenAI‑powered platform that now supports roughly 1,050 Japanese local governments and 550,000 public employees, aiming to become a shared operating system for municipal work.

Supporting Thailand’s next generation of AI startupsStartups

Supporting Thailand’s next generation of AI startups

OpenAI has partnered with Thailand's Ministry of Higher Education, Science, Research and Innovation to launch an eight-week accelerator program for ten local startups in the health, wellness, and education sectors.