Meta Faces Lawsuit Over Alleged Copyright Infringement in AI Training
Meta is alleged to have illegally copied millions of copyrighted works to train Llama, prompting a lawsuit from five publishers and author Scott Turow.
Examine your AI training data sources for potential copyright infringement and implement a compliance audit.
Summary
Meta is alleged to have illegally copied millions of copyrighted books, articles, and other works to train its Llama 1 model, prompting a lawsuit filed on May 5, 2026 by five major publishers—Hachette, Macmillan, McGraw Hill, Elsevier, and Cengage—alongside author Scott Turow.
The complaint claims that Meta, under Mark Zuckerberg’s direction, torrented 267 TB of pirated material, including content from LibGen, and that the company considered licensing up to $200 million for training data before abruptly halting the strategy. Meta also signed four licenses in 2022 with African-language publishers and agreements with Fox News, CNN, and USA Today, yet the lawsuit argues that Zuckerberg personally authorized and encouraged the infringement.
The suit seeks unspecified monetary damages and alleges that Llama’s outputs include verbatim copies, near‑verbatim reproductions, and derivative works that mirror the expressive elements of the original authors. Meta’s spokesperson maintains that training on copyrighted material can qualify as fair use, but the court has previously ruled in favor of fair use in similar cases.
The case highlights the growing legal scrutiny around AI training data and the need for companies to audit their data sources for compliance.
Key changes
- Meta allegedly copied millions of copyrighted books, articles, and other works to train Llama 1
- Five publishers—Hachette, Macmillan, McGraw Hill, Elsevier, Cengage—filed the lawsuit on May 5 2026
- Meta torrented 267 TB of pirated material, including LibGen content
- Meta considered licensing up to $200 million for training data before stopping the strategy
- Meta signed four licenses in 2022 with African-language publishers and agreements with Fox News, CNN, USA Today
- The lawsuit claims Zuckerberg personally authorized and encouraged the infringement
- Llama’s outputs include verbatim copies, near‑verbatim reproductions, and derivative works
- Meta’s spokesperson argues that training on copyrighted material can qualify as fair use