Phone Phone (Hover)
WhatsApp WhatsApp (Hover)
Phone
Call
++1(970)567-7400
WhatsApp
Whatsapp
Login In Sign up

Asia

North America

Asia

North America

Case Analysis: U.S. Copyright Infringement Dispute over Pirated Literary Works for AI Model Training

IPcrossark
Copyright
2026-08-10 07:18:16
 

 

This is a landmark class‑action case heard by the United States District Court for the Northern District of California. All corporate names are replaced with pseudonyms to protect commercial confidentiality. The plaintiff group consists of multiple published fiction and non‑fiction writers collectively named Global Author Collective. Their creative books were legally published and protected under U.S. Copyright Act Title 17. The defendant, Nova AI Labs, a fast‑growing artificial‑intelligence enterprise, developed a well‑known large‑language conversational model for global commercial subscription services. The company earned substantial revenue from paid subscriptions, enterprise licensing and customized model fine‑tuning services.

 

To build its core training dataset, Nova AI Labs adopted two different data‑collection channels. <b>One channel involved legally purchasing large volumes of physical printed books and scanning page‑by‑page for model training; the second channel downloaded roughly seven million copyrighted e‑books from unauthorised shadow‑library pirate websites without obtaining permission or paying licensing fees to original authors</b>. The pirated digital copies were permanently stored within Nova AI Labs internal cloud library and integrated into core training corpus for its generative AI model. The writers’ copyrighted novels, essays and biographical works were included within these pirated resources. After discovering this large‑scale unauthorized reproduction, several authors collected public technical disclosure documents, dataset metadata and related media reports, and filed a class‑action copyright infringement lawsuit on behalf of affected creators.

The plaintiffs alleged that Nova AI Labs committed direct copyright infringement through reproduction and storage of pirated literary works, and further used these copies for commercial model training. The plaintiffs requested declaratory judgment of copyright infringement, injunctive relief prohibiting further use of unlicensed pirated content, and massive monetary compensation for collective author losses. During litigation, Nova AI Labs raised the core fair‑use defence under Section 107 of the United States Copyright Act. <b>The defendant insisted that feeding copyrighted literary texts into large‑language‑model training constituted highly transformative use, analogous to human reading and learning; therefore, all reproduction activities should qualify for fair‑use exemption regardless of data source</b>. The AI developer emphasized that model training altered original expressive content into numerical weights rather than reproducing complete original texts in final outputs.

 

The federal district judge strictly applied the traditional four‑factor fair‑use test established by U.S. copyright precedent. <b>The four statutory fair‑use factors cover: purpose and character of use, nature of copyrighted original work, amount and substantiality of materials taken, and market harm caused to rights‑holders</b>. The court made a critical factual distinction between two separate data‑acquisition modes operated by Nova AI Labs. For physical books lawfully purchased and scanned by the defendant, the court held such training behaviour was highly transformative and could satisfy fair‑use requirements. However, the judge reached a completely different conclusion for e‑books downloaded from pirate shadow‑library platforms.

 

<b>Making permanent digital copies from pirated online sources bypasses normal publishing licensing markets and infers direct market substitution harm against authors; such behaviour cannot be classified as transformative fair‑use under U.S. copyright law</b>. Even though model training converts source text into neural‑network parameters, the initial act of mass downloading and permanent storage of full‑text pirated copyrighted works constitutes direct copyright reproduction infringement. The court pointed out that fair‑use protection does not immunise commercial companies from liability when they obtain source materials through infringing pirate channels. The source origin of training data carries decisive legal weight for fair‑use evaluation.

 

After multiple rounds of court‑ordered mediation, both parties reached a class‑action settlement agreement which received final judicial approval. <b>Nova AI Labs agreed to pay total settlement compensation of 1.5 billion US dollars to the affected author class, and must permanently delete all pirated e‑book files stored within its internal training library</b>. The defendant also promised to implement standardized data‑source review mechanisms for future dataset construction, avoiding materials obtained from known pirate shadow‑library websites. The settlement did not eliminate authors’ future statutory rights for new infringement disputes.

 

This ruling delivers far‑reaching compliance guidance for global generative‑AI enterprises operating inside the United States. <b>Transformative training purpose alone cannot guarantee fair‑use protection; AI developers must carefully verify the legal source of every training‑data input</b>. Lawful acquisition channels do not automatically mean all training activities are fair‑use; meanwhile, using obviously pirated copyrighted resources will substantially increase copyright‑infringement risks. Technology companies should establish copyright compliance workflows for dataset sourcing, maintain documentary evidence of legal data origin, and negotiate formal licensing agreements with publishers and author groups where commercially feasible.

 

Reference Links

 

1.IPcrossark:https://www.ipcrossark.com/en/trademark.html?cid=60 

2.https://www.copyright.gov/ai/ (U.S. Copyright Office official AI‑copyright policy portal)

3.https://www.courtlistener.com/docket/69058235/bartz-v-anthropic-pbc/ (Federal court case docket for AI book‑copyright class‑action litigation)

4. https://fairuse.stanford.edu/case/page/15/ (Stanford University Fair‑Use Case Database)