Sign in
Edition of 06:00 CETSunday, August 9, 2026
320 outlets · 17 languages454 briefings today
Wednesday, May 6, 2026

Publishers and Authors Sue Meta and Zuckerberg Over AI Training on Pirated Books

Five major publishing houses and novelist Scott Turow accuse Meta of illegally using millions of copyrighted works to train its Llama AI, with Zuckerberg’s alleged approval.

A class-action lawsuit filed in Manhattan federal court this week has brought the simmering conflict between creative industries and big tech to a new boil. Five of the world’s largest publishing houses — Elsevier, Cengage, Hachette, Macmillan, and McGraw Hill — together with bestselling author Scott Turow, accuse Meta and its chief executive, Mark Zuckerberg, of systematically pirating millions of copyrighted books and journal articles to train the company’s Llama family of language models. The complaint alleges that Meta’s engineers downloaded unlicensed copies from notorious pirate repositories such as Library Genesis and Anna’s Archive, and that Zuckerberg personally authorised the practice, invoking his company’s old motto of moving fast and breaking things.

This is far from an isolated skirmish. Viewed from Washington, the lawsuit fits into a broader pattern of legal challenges against Silicon Valley’s appetite for copyrighted training data. Authors including James Patterson and Donna Tartt have lent their names to earlier actions, and Meta already reached a $1.5 billion settlement with AI rival Anthropic in 2025 over similar allegations. In Berlin and Frankfurt, where publishers have long lobbied for stronger digital copyright protections, the case is seen as a pivotal test of whether “fair use” doctrines in the United States can be stretched to cover the mass ingestion of creative works without compensation. Meanwhile, legal observers in London note that the British government is closely watching the outcome, with its own AI white paper struggling to balance innovation incentives against creators’ rights.

Meta has promised to defend itself vigorously, arguing that training artificial intelligence on copyrighted material constitutes fair use — a position that has already won some support in lower US courts. But the sheer scale of the alleged infringement — the publishers claim Meta copied entire libraries, from academic textbooks to science fiction trilogies — narrows the company’s room for manoeuvre. The complaint explicitly seeks to certify a class of all affected rights holders, which, if granted, could expose Meta to damages running into billions of dollars.

What makes this litigation especially significant is its timing. With AI regulation still patchy across jurisdictions, the judiciary is increasingly being called upon to draw the boundaries of permissible training data use. From Tel Aviv to Sydney, analysts point out that a ruling against Meta would send shockwaves through the industry, forcing every developer of large language models to re-examine their data-sourcing practices. The case is widely expected to climb all the way to the Supreme Court, where it will ultimately decide whether the digital age’s most transformative technology can be built on a foundation of stolen words.

Breaking
Iran-Oman Hormuz shipping deal nears, but Tehran ties full reopening to US concessions·Mild earthquake and tremors cause cracks in houses in Maharashtra’s Nashik district·Top US general privately urges exit strategy from Iran war, warning air power alone will fail·Berkshire Hathaway ends three-year selling streak as Abel puts cash to work·Tirante advances in Montreal as upsets leave Shelton lone top-10 survivor·IAF officer charged over alleged honey-trap leak of defence secrets·One Night Only director cut street-shot nudity after test audiences squirmed·Zebra-striped cows halve biting-fly landings, strengthening insect-deterrence hypothesis·Iran-Oman Hormuz shipping deal nears, but Tehran ties full reopening to US concessions·Mild earthquake and tremors cause cracks in houses in Maharashtra’s Nashik district·Top US general privately urges exit strategy from Iran war, warning air power alone will fail·Berkshire Hathaway ends three-year selling streak as Abel puts cash to work·Tirante advances in Montreal as upsets leave Shelton lone top-10 survivor·IAF officer charged over alleged honey-trap leak of defence secrets·One Night Only director cut street-shot nudity after test audiences squirmed·Zebra-striped cows halve biting-fly landings, strengthening insect-deterrence hypothesis·
Upd. 07:13 PM5 languages · 8 outlets
8 outlets|5 languages|3 min read
Wednesday, May 6, 2026

Publishers and Authors Sue Meta and Zuckerberg Over AI Training on Pirated Books

Five major publishing houses and novelist Scott Turow accuse Meta of illegally using millions of copyrighted works to train its Llama AI, with Zuckerberg’s alleged approval.

A class-action lawsuit filed in Manhattan federal court this week has brought the simmering conflict between creative industries and big tech to a new boil. Five of the world’s largest publishing houses — Elsevier, Cengage, Hachette, Macmillan, and McGraw Hill — together with bestselling author Scott Turow, accuse Meta and its chief executive, Mark Zuckerberg, of systematically pirating millions of copyrighted books and journal articles to train the company’s Llama family of language models. The complaint alleges that Meta’s engineers downloaded unlicensed copies from notorious pirate repositories such as Library Genesis and Anna’s Archive, and that Zuckerberg personally authorised the practice, invoking his company’s old motto of moving fast and breaking things.

This is far from an isolated skirmish. Viewed from Washington, the lawsuit fits into a broader pattern of legal challenges against Silicon Valley’s appetite for copyrighted training data. Authors including James Patterson and Donna Tartt have lent their names to earlier actions, and Meta already reached a $1.5 billion settlement with AI rival Anthropic in 2025 over similar allegations. In Berlin and Frankfurt, where publishers have long lobbied for stronger digital copyright protections, the case is seen as a pivotal test of whether “fair use” doctrines in the United States can be stretched to cover the mass ingestion of creative works without compensation. Meanwhile, legal observers in London note that the British government is closely watching the outcome, with its own AI white paper struggling to balance innovation incentives against creators’ rights.

Meta has promised to defend itself vigorously, arguing that training artificial intelligence on copyrighted material constitutes fair use — a position that has already won some support in lower US courts. But the sheer scale of the alleged infringement — the publishers claim Meta copied entire libraries, from academic textbooks to science fiction trilogies — narrows the company’s room for manoeuvre. The complaint explicitly seeks to certify a class of all affected rights holders, which, if granted, could expose Meta to damages running into billions of dollars.

What makes this litigation especially significant is its timing. With AI regulation still patchy across jurisdictions, the judiciary is increasingly being called upon to draw the boundaries of permissible training data use. From Tel Aviv to Sydney, analysts point out that a ruling against Meta would send shockwaves through the industry, forcing every developer of large language models to re-examine their data-sourcing practices. The case is widely expected to climb all the way to the Supreme Court, where it will ultimately decide whether the digital age’s most transformative technology can be built on a foundation of stolen words.

Source divergence

— · 8 outlets · 5 languages

0%Low

How sources tell the same facts differently.

This story appeared in

8 outlets · 5 languages

Broaden your view

From Geopolitics & Politics

US Senate votes 86-11 to advance Russia sanctions bill authorising 100% tariffs on top energy buyers

2 languages · 40 outlets

From Economy & Markets

US imposes 15% tariff and price floors on polysilicon to counter China’s supply-chain dominance

4 languages · 16 outlets

From Technology

India cuts AI-content takedown deadline to three hours after Meta row

2 languages · 8 outlets

Read more