Rows of hardcover law books on library shelves sealed behind a locked wrought-iron gate, representing restricted and licensed content in the AI training data lawsuits.
All Posts

AI Training Data Lawsuits: What They Mean for Web Scrapers

Ryan Turner
Ryan Turner · Head of Innovation
Open markdown

In September 2025, a federal judge granted preliminary approval to a $1.5 billion settlement, the largest copyright payout in AI history. The claims behind it: Anthropic trained its models on roughly 500,000 pirated books (Copyright Alliance, "Participating in the Bartz v. Anthropic Settlement," retrieved 2026-09-17). That figure alone should get the attention of anyone who collects data for a living.

Over the past eighteen months, courts on two continents have started drawing real lines around how AI companies and scrapers can acquire training data, and those lines don't always match what the industry assumed. This piece walks through the major rulings, what they actually decided, and what data teams and scraping infrastructure providers should change about how they operate. For background on the mechanics, see how web scraping works.

Key Takeaways
  • Anthropic's $1.5B settlement shows piracy, not training, created the liability (Copyright Alliance, 2026).
  • Ross Intelligence lost its fair-use defense because its AI tool competed directly with Westlaw's own market; Anthropic and Meta never had to answer that question.
  • Just 14% of top domains have set an AI-bot robots.txt signal (Cloudflare Radar, 2025).
  • Evasion gets you blacklisted. No lawsuit needed.

What Are Courts Actually Ruling on AI Training Data?

Courts are ruling that how you acquire training data can create liability entirely separate from how you use it. The Bartz v. Anthropic settlement broke down to roughly $3,000 per work across about 500,000 pirated books, a number tied directly to acquisition method, not to how the books were later used in training (Copyright Alliance, "Participating in the Bartz v. Anthropic Settlement," retrieved 2026-09-17).

Judge William Alsup's underlying ruling in Bartz v. Anthropic drew a sharp line the industry didn't fully expect. Training on books Anthropic had purchased and scanned was fair use, "spectacularly transformative" in his words (Norton Rose Fulbright, Inside Tech Law, "Bartz v. Anthropic: Settlement reached after landmark summary judgment," retrieved 2026-09-17). Downloading pirated copies to build a permanent research library was a separate act, and it wasn't fair use no matter how the data was later used.

That distinction matters more than the settlement number itself. Acquisition and use are now two different legal questions, and a company can win on one while losing badly on the other. Anthropic reached its settlement agreement in August 2025, and Judge Alsup granted preliminary approval the following month (CNBC, "Judge grants preliminary OK to $1.5B settlement with authors," Sept 25 2025, retrieved 2026-09-17).

Is a well-funded AI lab really at less legal risk than a small scraper using the same pirated source? Not necessarily. The settlement shows scale doesn't shield anyone from an acquisition-method claim once it's established as independent from the fair-use question. See what counts as AI training data for the underlying mechanics these rulings are judging.

The 2025 Pivot From Litigation to Licensing Horizontal timeline showing four 2025 AI copyright settlement events: Anthropic agrees to a 1.5 billion dollar settlement in Bartz v. Anthropic (August 2025); the court grants preliminary approval (September 2025); Universal Music Group settles with Udio and launches a licensed AI music platform (October 2025); Warner Music Group settles with both Udio and Suno (November 2025). The 2025 Pivot From Litigation to Licensing Major AI copyright settlements, Aug-Nov 2025 $1.5B settlement Bartz v. Anthropic Aug 2025 Court approval Preliminary approval granted Sep 2025 UMG x Udio deal Licensed AI music platform Oct 2025 WMG settles both Udio, then Suno Nov 2025 Source: Copyright Alliance; CNBC; Music Business Worldwide (2025)
The 2025 pivot from AI copyright litigation to settlement-plus-licensing (Source: Copyright Alliance, CNBC, Music Business Worldwide, 2025)

Why Did Thomson Reuters Beat Ross Intelligence When Anthropic Won on Fair Use?

Thomson Reuters won because Ross Intelligence built a directly competing product from licensed content it wasn't authorized to use for training. In February 2025, Judge Stephanos Bibas (D. Del.) granted partial summary judgment to Thomson Reuters, the first summary-judgment loss for an AI company on a training-data fair use defense (Goodwin Procter, "Court Rejects Fair Use Defense in AI Copyright Case," retrieved 2026-09-17).

Ross had used Westlaw headnotes, editorial content Thomson Reuters spent decades building, to train a legal-research tool that competed head-on with Westlaw itself. Judge Bibas found the use non-transformative and commercially damaging to the market for the original work (Jenner & Block, client alert, retrieved 2026-09-17).

Compare that to Kadrey v. Meta. In June 2025, Judge Vince Chhabria (N.D. Cal.) granted Meta partial summary judgment on fair use, even though some training books came from shadow libraries, because plaintiffs couldn't show Meta's use diluted the market for their work (Goodwin Procter, retrieved 2026-09-17). He was careful to note a future plaintiff who could prove market dilution might win where these plaintiffs didn't (Copyright Alliance, "Kadrey v. Meta Decision," retrieved 2026-09-17).

So what actually separates a winning fair use defense from a losing one? Market competition, it turns out, not the mechanics of the copying itself.

Read together, these rulings suggest acquisition method and competitive use function as two independent liability tracks. Anthropic lost on acquisition (piracy) despite winning on transformative use. Ross lost on competitive use despite drawing from licensed source content. Neither factor guarantees protection on its own.

Case What was trained on How the data was acquired Competes with the source? Fair use outcome
Bartz v. Anthropic Books, purchased and scanned Legally purchased copies (lawful) plus pirated copies from shadow libraries (unlawful) No Training on purchased books: fair use. Downloading pirated copies: not fair use, leading to the $1.5B settlement
Kadrey v. Meta Books, including from shadow libraries Partly pirated Plaintiffs did not prove market harm Fair use on this record; judge left the door open for plaintiffs who can show market dilution
Thomson Reuters v. Ross Intelligence Westlaw headnotes Licensed content, used without authorization for training Yes, a competing legal-research product Not fair use, the first summary judgment loss on AI training
Getty Images v. Stability AI (UK) Licensed stock images Used without a license; training occurred outside the UK Yes, a competing image-licensing business Copyright claim rejected on jurisdictional grounds; trademark infringement found instead

The Cases Still Working Their Way Through Court

Several of the biggest AI copyright cases remain undecided, and the outcomes could reshape fair-use doctrine further. The consolidated NYT and Authors Guild suits against OpenAI are now In re OpenAI Copyright Litigation, before Judge Sidney Stein. The case is in active discovery as of September 2026, with no trial date set (Washington Post, "DOJ urges judge to rule for OpenAI, Microsoft in NY Times lawsuit," Sept 2 2026, retrieved 2026-09-17).

Judge Stein largely denied OpenAI and Microsoft's motion to dismiss in April 2025, letting most claims proceed and explicitly declining to call AI training "inherently transformative." An October 27, 2025 order went further: Judge Stein held that ChatGPT-generated plot summaries of plaintiffs' novels could themselves constitute infringement, an "outputs" theory distinct from the training question (IPWatchdog, "OpenAI Loses Bid to Dismiss Multi-District Class Action Over ChatGPT Outputs," Oct 28 2025, retrieved 2026-09-17).

The case has drawn outside attention too. On September 2, 2026, the Department of Justice filed a brief urging the court to rule for OpenAI and Microsoft (Washington Post, retrieved 2026-09-17; Axios, "NYT, OpenAI, Microsoft copyright lawsuit," Sept 8 2026, retrieved 2026-09-17).

Getty Images v. Stability AI split along an ocean. On November 4, 2025, the UK High Court ruled that Getty's copyright claim failed because Stability's training occurred outside the UK, but found trademark infringement because early Stable Diffusion outputs carried the Getty watermark (Ropes & Gray, retrieved 2026-09-17). Getty's separate US case, refiled in N.D. Cal., survived a motion to dismiss on trademark and dilution claims in April 2026, with trial set for January-February 2028 (CourtListener docket, retrieved 2026-09-17).

The same theory, that AI outputs carrying a brand's watermark or recognizable derivative content amount to infringement, is now live in UK and US courts simultaneously. Copyright claims are getting narrowed on jurisdictional and transformative-use grounds, but trademark and output-based claims are filling the gap those narrower rulings leave behind.

The pattern extends beyond text. Disney, Universal, and Warner Bros. sued Midjourney between June and September 2025 over image generation, and the consolidated cases remain in a discovery dispute as of mid-2026 (Variety, retrieved 2026-09-17; IPWatchdog, "Warner Bros. Complaint Alleges Midjourney's Copyright Infringement Is 'Systematic,' 'Willful,'" retrieved 2026-09-17). In music, Universal Music Group settled with Udio in October 2025, pairing compensation with a licensing deal for a new AI music platform, and Warner Music settled with both Udio and Suno in November 2025 (Music Business Worldwide, retrieved 2026-09-17). Sony remains in litigation with both platforms. The pressure isn't limited to courtrooms either: see why access to web data is getting harder for the infrastructure side of the same squeeze.

Does Any of This Apply to Scraping Public Data, Not Just AI Training?

Scraping publicly accessible web pages is not, by itself, a federal crime, but that doesn't mean it's risk-free. hiQ Labs v. LinkedIn settled that question at the Ninth Circuit: scraping public pages doesn't violate the CFAA's "without authorization" provision, though breach-of-contract, copyright, and trespass claims remain fully available (White & Case, "Web scraping, website terms and the CFAA," retrieved 2026-09-17).

That ruling came on remand after the Supreme Court's Van Buren decision narrowed the CFAA's scope. The same proceedings also found hiQ still breached LinkedIn's Terms of Service, a separate contract claim untouched by the CFAA holding (Fenwick & West, "hiQ Labs Scrapes By Again," retrieved 2026-09-17). In practice, the real risk question shifted from "is this hacking" to "did I have permission or a license."

If scraping public pages isn't hacking, why do companies still get blocked and sued over it? Because CFAA immunity never covered contract, copyright, or trespass claims, and that's exactly where most disputes now land.

A separate incident in August 2025 showed how fast informal enforcement can move. Cloudflare reported that Perplexity used undeclared, rotating user-agents and ASNs to keep crawling test domains that had explicitly disallowed all crawlers via robots.txt (Search Engine Journal, "Cloudflare Delists And Blocks Perplexity From Crawling Websites," retrieved 2026-09-17). Cloudflare delisted Perplexity's verified-bot status and blocked its crawler network-wide, no court filing required.

Perplexity publicly disputed Cloudflare's characterization of events (Search Engine Journal, retrieved 2026-09-17). Regardless of how that dispute settles, the lesson holds: getting caught evading a declared block can end a crawler's access to large parts of the web overnight. For a closer look at where the legal lines sit, see ethical web scraping practices.

How Infrastructure Providers Are Responding

Infrastructure providers stopped waiting for courts and started changing defaults themselves. Starting July 1, 2025, Cloudflare made blocking unauthorized AI crawlers the default setting for every new domain, flipping the prior opt-in-block model to opt-out-access (Cloudflare, press release, "Cloudflare just changed how AI crawlers scrape the internet, at large," retrieved 2026-09-17).

Cloudflare paired that shift with "Pay Per Crawl," a mechanism letting site owners charge crawlers for access, which the company says will evolve into "Pay Per Use" starting September 15, 2026, blocking mixed-purpose crawlers by default on ad-supported pages (Cloudflare Blog, retrieved 2026-09-17). In a traffic window covering June 19-26, 2025, Cloudflare Radar found that Anthropic's Claude made roughly 71,000 HTML page requests for every one referral it sent back to the sites it crawled (Cloudflare Blog/Radar, "AI crawl-to-referral ratio on Radar," retrieved 2026-09-17).

Most Sites Still Have No AI-Bot Signal Set Donut chart showing that of the top 10,000 web domains with a robots.txt file, only 14 percent have any AI-bot-specific directive; 86 percent have no AI-bot signal. Source: Cloudflare Radar, June 2025. Most Sites Still Have No AI-Bot Signal Set Share of top 10,000 domains with an AI-bot robots.txt directive 14% have an AI-bot robots.txt signal 14%: set an AI-bot-specific robots.txt directive 86%: no AI-bot signal set Source: Cloudflare Radar (June 2025)
Only 14% of the top 10,000 domains have set an AI-bot-specific robots.txt directive (Source: Cloudflare Radar, June 2025)

Most site owners still haven't set any signal at all. GPTBot was the single most-blocked crawler in that same dataset, disallowed by 312 of 3,816 domains with a robots.txt file, against just 61 that explicitly allowed it (Cloudflare Blog/Radar, retrieved 2026-09-17).

Regulators are starting to treat machine-readable signals as the only ones that count. On December 10, 2025, the Hanseatic Higher Regional Court of Hamburg ruled in Kneschke v. LAION (OLG Hamburg, 5 U 104/24) that a text-and-data-mining opt-out is legally effective only if it's machine-readable. Robots.txt, an X-Robots-Tag header, or the TDM Reservation Protocol all count. A clause buried in natural-language website terms does not (Norton Rose Fulbright, Inside Tech Law, "Machine-Readable Opt-Outs and AI Training: Hamburg Court Clarifies Copyright Exceptions," retrieved 2026-09-17).

That approach lines up with the EU's GPAI Code of Practice, published in July 2025 and signed by 23 AI developers by July 2026, including OpenAI, Google, Anthropic, and Microsoft, all of whom committed to crawling with robots.txt-aware bots and honoring machine-readable reservations.

Infrastructure providers and regulators are effectively out-pacing the courts here. Cloudflare flipped its default crawler policy in July 2025, and the Hamburg court's machine-readable-signal standard arrived well before US case law settled comparable questions about scraping and licensing. Teams waiting for litigation to resolve before adjusting policy may find the defaults have already changed underneath them.

What Scrapers and Data Teams Should Do Now

Data teams that treat provenance as an afterthought are the ones most exposed to the risks these rulings created. The Bartz settlement alone cost roughly $3,000 per pirated work across about 500,000 books (Copyright Alliance, retrieved 2026-09-17), a number any legal team can multiply against its own unlicensed archive.

1. Track data provenance end to end. Know exactly where every training or collection input came from, and whether it was licensed, purchased, or scraped. The Anthropic case shows regulators and courts will separate acquisition method from downstream use when assigning liability.

2. Respect machine-readable opt-out signals. robots.txt, ai.txt, and the TDM Reservation Protocol are the standards regulators now recognize, even where a jurisdiction hasn't forced the issue yet.

3. Favor licensing over piracy-adjacent shortcuts. Purchased or licensed data relationships hold up in court in a way pirated archives don't, since Anthropic's fair-use win on purchased books did nothing to offset its liability for the pirated copies sitting next to them in the same training run.

4. Avoid identity-evasion tactics. Rotating user-agents to dodge a declared block gets you blacklisted. No court filing required. Background on what AI scraping actually involves is a useful starting point for auditing your own pipeline.

Provenance is one piece of a defensible data-sourcing posture, not the whole picture. Massive's residential proxy network runs on real user devices across 195+ countries, every device opted in through the Massive SDK, delivering clean HTML or markdown from any public source in any location. It's SOC 2 audited and GDPR compliant, one input into a defensible posture, not a legal shield or a substitute for legal advice.

Conclusion

The legal picture for AI training data split into two separate tracks this year: how you acquire data, and how you use it, and courts are now willing to punish failures on either one independently.

  • Acquisition method is now its own liability track, separate from fair use.
  • Competitive use kills a fair-use defense faster than piracy does: Ross Intelligence lost outright because its tool competed directly with Westlaw, a test Anthropic and Meta never had to face.
  • Infrastructure is moving faster than courts. Cloudflare flipped its crawler default in July 2025; EU regulators followed months later.

Expect more settlement-plus-licensing deals like UMG's and Warner Music's through 2027, as the "sue then license" pattern keeps spreading across text, image, and music. For the practical side of sourcing defensibly, see building a data pipeline on live web data.

Sources

  1. Copyright Alliance, "Participating in the Bartz v. Anthropic Settlement," retrieved 2026-09-17
  2. CNBC, "Judge grants preliminary OK to $1.5B settlement with authors," retrieved 2026-09-17
  3. Norton Rose Fulbright / Inside Tech Law, "Bartz v. Anthropic: Settlement reached after landmark summary judgment and class certification," retrieved 2026-09-17
  4. Goodwin Procter, "Court Rejects Fair Use Defense in AI Copyright Case," retrieved 2026-09-17
  5. Jenner & Block, client alert on Thomson Reuters v. Ross Intelligence, retrieved 2026-09-17
  6. Goodwin Procter, "Northern District of California Judge Rules" (Kadrey v. Meta), retrieved 2026-09-17
  7. Copyright Alliance, "Kadrey v. Meta Decision," retrieved 2026-09-17
  8. Washington Post, "DOJ urges judge to rule for OpenAI, Microsoft in NY Times lawsuit," retrieved 2026-09-17
  9. Axios, "NYT, OpenAI, Microsoft copyright lawsuit," retrieved 2026-09-17
  10. IPWatchdog, "OpenAI Loses Bid to Dismiss Multi-District Class Action Over ChatGPT Outputs," retrieved 2026-09-17
  11. Ropes & Gray, "Getty Image Loses Copyright Infringement Claim Against Stability AI in UK's First..." retrieved 2026-09-17
  12. Cleary Gottlieb, "UK High Court Issues Landmark Ruling in Getty Images v. Stability AI," retrieved 2026-09-17
  13. CourtListener, docket for Getty Images (US) Inc. v. Stability AI Ltd., retrieved 2026-09-17
  14. Music Business Worldwide, "Universal Music settles Udio lawsuit, strikes deal for licensed AI music platform," retrieved 2026-09-17
  15. Hollywood Reporter, "Universal Music Group Announces Settlement With Udio," retrieved 2026-09-17
  16. Variety, on the Midjourney studios discovery dispute, retrieved 2026-09-17
  17. IPWatchdog, "Warner Bros. Complaint Alleges Midjourney's Copyright Infringement Is 'Systematic,' 'Willful,'" retrieved 2026-09-17
  18. White & Case, "Web scraping, website terms and the CFAA: hiQ's preliminary injunction affirmed again," retrieved 2026-09-17
  19. Fenwick & West, "hiQ Labs Scrapes By Again: The Ninth Circuit Reaffirms That Data Scraping Does Not Violate the CFAA," retrieved 2026-09-17
  20. Cloudflare Blog/Radar, "AI crawl-to-referral ratio on Radar," retrieved 2026-09-17
  21. Cloudflare, press release, "Cloudflare just changed how AI crawlers scrape the internet, at large," retrieved 2026-09-17
  22. Search Engine Journal, "Cloudflare Delists And Blocks Perplexity From Crawling Websites," retrieved 2026-09-17
  23. Norton Rose Fulbright, Inside Tech Law, "Machine-Readable Opt-Outs and AI Training: Hamburg Court Clarifies Copyright Exceptions," retrieved 2026-09-17

Frequently Asked Questions

Is it illegal to scrape publicly available websites?+

Not inherently. The Ninth Circuit held in hiQ Labs v. LinkedIn that scraping publicly accessible pages doesn't violate the CFAA (White & Case, retrieved 2026-09-17). But breach-of-contract, copyright, and trespass-to-chattels claims remain fully available, so a scraper can still face real liability without breaking that one statute. See how to give AI agents live web access for the robots.txt and crawler-access side of this.

Does training an AI model on copyrighted material count as fair use?+

It depends on acquisition and competition, not just transformation. Judge Alsup ruled training on purchased books was fair use in the Anthropic case, while Judge Bibas ruled Ross Intelligence's training on Westlaw headnotes for a competing product was not, the first such loss (Goodwin Procter, retrieved 2026-09-17).

What is the difference between the Anthropic and Thomson Reuters AI copyright rulings?+

Anthropic won on transformative training but lost on piracy, leading to a $1.5 billion settlement over roughly 500,000 pirated books (Copyright Alliance, retrieved 2026-09-17). Ross Intelligence lost outright because it used licensed Westlaw content to build a directly competing product, an outcome about market competition rather than the training method itself.

How can data teams reduce legal risk when collecting training data?+

Track provenance for every input, honor machine-readable opt-out signals like robots.txt and TDM Reservation Protocol, and favor licensed or purchased data over piracy-adjacent shortcuts. Cloudflare found only about 14% of top domains had set any AI-bot robots.txt signal as of June 2025 (Cloudflare Radar, retrieved 2026-09-17), leaving most exposure unaddressed by default.