\n\n\n\n Reading Is Free, Shoplifting Is Not - AgntHQ \n

Reading Is Free, Shoplifting Is Not

📖 4 min read•749 words•Updated Aug 24, 2026

Imagine walking into a bookstore, buying a stack of novels, reading every one, and then writing your own book informed by everything you absorbed. Nobody calls the cops. Now imagine climbing through the bookstore’s back window at 2am, walking out with the same stack, and reading them just as carefully. The reading is still fine. The window is the problem.

That, roughly, is where U.S. courts have landed on training AI models on copyrighted books. The learning is legal. The acquisition might not be. And the AI industry spent years hoping nobody would notice the difference.

Two Cases, One Uncomfortable Split

Two decisions out of the Northern District of California — Kadrey v. Meta Platforms and Bartz v. Anthropic PBC — did the work of separating those two questions. In Bartz, the court found that training AI on copyrighted works was fair use. It also denied summary judgment for Anthropic on its use of pirated copies to assemble a central library. Same case, two very different answers.

A federal court has described AI training on copyrighted books as “quintessentially” transformative fair use. That’s strong language. Transformative use is the load-bearing wall of fair use analysis, and courts do not reach for “quintessentially” casually.

Then there’s the number that actually got the industry’s attention. Judge William Alsup ordered Anthropic to pay a $1.5 billion copyright settlement to a group of writers whose works were used to train the company’s models. Not for the training. For how the books got there.

What This Actually Means for AI Companies

I review AI tools for a living, and I’ve watched a lot of founders treat data sourcing as a rounding error. Scrape it, torrent it, grab the shadow library, sort out the paperwork later. The reasoning was always the same: the models are transformative, so the inputs don’t matter.

Half right. Which, legally speaking, is the expensive half.

The practical takeaways from these rulings are narrower than the headlines suggest:

  • Legal acquisition matters independently of use. Buying, licensing, or otherwise lawfully obtaining works is its own question, separate from what you do with them afterward.
  • Transformative use is a real defense. Courts are treating model training as something meaningfully different from reproduction.
  • Piracy does not get laundered by transformation. A downstream fair use finding does not retroactively bless how the material was collected.
  • The exposure is enormous. A billion-plus dollar settlement is not a cost of doing business for most companies. It’s an extinction event.

Why “It’s Complicated” Is the Honest Answer

I don’t like hedging. But anyone giving you a clean yes or no on this is selling something.

These are district court decisions. They’re persuasive, they’re being cited, and they’re shaping how legal teams advise clients right now. They are not the last word. A March 2026 publication from Neal, Gerber & Eisenberg framing recent decisions as raising critical concerns tells you the legal profession isn’t treating this as settled either.

What we have is a workable framework, not a finished rulebook: acquire lawfully, use transformatively, and keep watching the courts. That’s genuinely useful guidance. It’s also the kind of guidance that shifts when an appellate court takes an interest.

The Part Nobody Wants to Hear

If you’re building on top of AI models — agents, wrappers, fine-tuned products, whatever — this isn’t purely somebody else’s legal problem. You’re building on a foundation whose provenance you probably can’t audit. Most model providers won’t tell you exactly what went into training. Some of them may not be able to reconstruct it.

That’s not a reason to stop building. It is a reason to ask vendors uncomfortable questions and to notice which ones answer with specifics versus which ones answer with vibes. The companies that invested in licensing and documented sourcing early are about to look a lot smarter than the ones that treated data acquisition as a scraping exercise.

I’ve been skeptical of the “AI training is theft” framing for a while, and these rulings support that skepticism on the narrow question of learning from text. Reading isn’t stealing, and pattern extraction from lawfully obtained books looks a lot like reading at scale. But I’ve also been skeptical of the industry’s assumption that scale excuses sourcing, and the $1.5 billion settlement suggests courts share that skepticism.

The distinction courts are drawing is not complicated to understand. It’s just inconvenient for anyone who built their pipeline before the rules got clarified. Legal front door, transformative use, defensible position. Back window, and it doesn’t matter how clever the model turned out.

🕒 Published:

📊
Written by Jake Chen

AI technology analyst covering agent platforms since 2021. Tested 40+ agent frameworks. Regular contributor to AI industry publications.

Learn more →
Browse Topics: Advanced AI Agents | Advanced Techniques | AI Agent Basics | AI Agent Tools | AI Agent Tutorials
Scroll to Top