Copyright law struggles to keep pace with AI training on books
Reported by TechCrunch AI: Is it legal to train AI models on copyrighted books? It’s complicated. Analysis and context written by TickrWire.
Courts are split on whether training AI on copyrighted books counts as illegal copying or protected fair use, leaving authors and tech firms in legal limbo.

- Judge William Alsup ruled that training AI on copyrighted books is lawful, comparing it to reading rather than copying, but Anthropic was fined $1.5 billion for using illegal shadow libraries to obtain the books.
- Courts are split on whether AI training constitutes fair use, with some rulings favoring transformative use and others penalizing direct competition with original works.
- The Thaler v. Perlmutter case established that AI-generated works cannot be copyrighted, complicating attribution and ownership in AI-assisted content creation.
- Legal uncertainty is forcing AI companies to either license content or avoid high-risk training materials, while authors face potential market erosion from AI-generated content.
- Without a unified legal framework, the industry must adapt to a patchwork of rulings that could be overturned by higher courts or new legislation.
The rapid rise of large language models has collided with a fundamental question in copyright law: when an AI system ingests millions of books to learn patterns, is that training legal or does it infringe on authors’ rights? The answer, according to recent court rulings and legal experts, is far from settled. While some judges have ruled that AI training resembles reading a book rather than copying it, others have drawn sharp distinctions between transformative use and direct competition, leaving the field in a state of legal uncertainty that affects everyone from authors to AI developers.
In one of the first major cases to test these boundaries, Judge William Alsup of the Northern District of California ruled last year that Anthropic’s use of copyrighted books to train its AI models was lawful. The judge compared the process to a writer studying literature, emphasizing that the AI was not replicating or supplanting the original works but using them to create something new. The ruling came despite a $1.5 billion settlement Anthropic paid to a group of authors, but the penalty was tied to the company’s use of illegal shadow libraries to obtain the books, not the act of training itself. Legal observers like Cathy Gellis, an intellectual property attorney, argue that this distinction is critical. She points out that copyright law focuses on copying, not mere consumption or study of a work. "Copyright law hinges on copying, but it doesn’t hinge on using the work or experiencing the work," Gellis told TechCrunch. This interpretation suggests that AI training could fall under fair use, a legal doctrine that allows limited use of copyrighted material without permission for purposes like criticism, education, or transformation.
However, not all courts have adopted this view. In a separate case, Thomson Reuters sued Ross Intelligence for copying its content to build a competing AI-based legal research platform. Judge Stephanos Bibas ruled that Ross’s use was not transformative because it did not serve a different purpose or character than Reuters’ original content. This decision highlights a key tension in AI copyright cases: when AI training leads to a product that directly competes with the original work, courts are more likely to find it unlawful. While authors could argue that chatbots trained on their books compete with them by generating synthetic content, this argument has not yet succeeded in court. The lack of a clear precedent means that AI companies and content creators are operating in a legal gray area where outcomes can vary dramatically depending on the judge and jurisdiction.
The ambiguity extends beyond training to the ownership of AI-generated content. In the case of Thaler v. Perlmutter, a federal court ruled that a work entirely created by AI cannot be copyrighted, raising new questions about how to attribute creativity in an era where tools like spell check or AI assistants play a role in content creation. "If you write your novel in Microsoft Word and run spell check, we kind of feel comfortable with the idea of saying that Word does not own your novel," Gellis noted. "AI is forcing us to look at a whole bunch of decisions that we kind of ignored for a while." This ruling complicates efforts to determine whether a book, article, or other creative work was generated with AI assistance, let alone how to compensate the original authors whose works may have contributed to the training data.
Legal experts warn that the current patchwork of rulings is shaping the industry in real time, even as definitive answers remain elusive. "What you are seeing is that the initial opening volleys are being influential, and that influence itself could be undone if other courts decide different things," Gellis said. "It will take later stages of litigation to figure out which one will prevail." This uncertainty is forcing AI companies to navigate a minefield of potential lawsuits, with some opting to license content or avoid high-risk training materials altogether. Meanwhile, authors and other content creators are left grappling with the erosion of their livelihoods as AI systems trained on their work generate competing or derivative content without compensation.
The stakes are high for both sides. For AI developers, the cost of litigation and potential damages could stifle innovation, while for authors and publishers, the threat of AI-generated content undermining their market is existential. The lack of a unified legal framework means that companies and creators must adapt to a shifting landscape where each new ruling could redefine the rules. Until Congress or the Supreme Court provides clarity, the debate will continue to play out in courtrooms, boardrooms, and legislative halls, with no clear end in sight. In the meantime, the industry’s reliance on copyrighted material for training underscores the urgent need for a solution that balances innovation with fair compensation for creators.
AI developers must navigate evolving copyright laws to avoid litigation and ensure their training data is legally sourced.
Companies using AI systems trained on copyrighted material risk lawsuits and financial penalties, while content creators face potential loss of revenue to AI-generated competitors.
Understanding the legal landscape of AI and copyright is crucial for future technologists and legal professionals navigating this emerging field.
The debate over AI training on copyrighted books highlights broader tensions between innovation and intellectual property rights.
- fair use
- A legal doctrine that allows limited use of copyrighted material without permission for purposes like criticism, education, or transformation.
- shadow libraries
- Online repositories of copyrighted books and other materials distributed without authorization, often used to train AI models.
- transformative use
- A legal concept where the use of a copyrighted work adds new expression or meaning, rather than merely copying the original.
AI bias estimate: The source leans toward a legal perspective that may understate the ethical and economic concerns of authors whose works are used without consent. (Automated estimate, not a definitive judgement.)
SecurityFlock CEO calls for ‘compromise’ as surveillance company faces growing backlash
SecurityHow China's gray market sells Claude tokens at a fraction of the price
SecurityFrontier AI labs still won’t say how they’d contain a rogue model
SecurityPsychological methods reveal major weaknesses in AI security testing
New White House strategy clarifies military tech priorities: undersea, outer space and AI - Breaking Defense
HardwareCerebras unveils CS-4 with double the performance on the same chip
Cerebras has launched the CS-4, a rack-scale AI accelerator that doubles performance over its predecessor by optimizing power and cooling for the WSE-3 chip.
AI ResearchWho’s behind the new ‘stealth model’ Ox Alpha?
A mysterious reasoning model named Ox Alpha appeared on OpenRouter, prompting widespread speculation regarding its anonymous creator.
AI ToolsAn AI boss fired its first employee but only after humans reminded it of its own rules
An AI agent running a San Francisco store fired an employee only after humans reminded it of its own termination rules, highlighting gaps in long-term memory and leniency in AI management.
AI ResearchAI could make scientists do more work less well, not less work better, study argues
A theoretical economics study argues that language models might make scientific research shallower because time saved on routine tasks encourages academics to start more projects rather than improve existing ones.
AI ToolsVercel Introduces ‘Is Agentic’, a Free Agent-Readiness Scoring Tool That Audits Public Websites Using Ora’s 100+ Checks
Vercel and Ora launch Is Agentic, a free tool that scores how easily AI agents can discover, access, understand, and use a website using over 100 checks across four layers.
BusinessHarvard’s $699 startup bootcamp offers AI avatars of its instructors
Harvard Business School’s eight‑week Foundry bootcamp now includes AI avatars from HeyGen that give feedback on practice pitches and board meetings, at a price of $699.