Training an AI model on a copyrighted novel is legal. Downloading that same novel from a piracy site to do it is not. Same book, same model, same commercial product at the end of the pipeline, and the difference between a dismissed lawsuit and a very expensive trial comes down to whether someone paid for the copy.
That is where US courts have landed, and it is stranger than either side of the AI debate wants to admit.
What the courts actually said
Two decisions out of the Northern District of California set the tone. In Bartz v. Anthropic PBC, Judge Alsup ruled that training large language models on copyrighted books counts as fair use. Authors had sued Anthropic over exactly this. The court granted summary judgment on the training question. A federal court has described the act as “quintessentially” transformative fair use, which is about as strong a signal as courts give.
In Kadrey v. Meta Platforms, Inc., Judge Chhabria reached a similar conclusion about Meta’s training on copyrighted works. Two judges, two cases, same direction of travel on the core question.
Then comes the part everyone skips when they quote these rulings on social media. The fair use finding covered training on legally acquired material. The claims involving pirated copies were not resolved in the AI companies’ favor. Those are heading toward trial, where they will be decided separately.
So the legal picture is not “AI training is legal.” It is closer to: the transformation is fine, the acquisition might not be.
Why this split makes more sense than it sounds
My first reaction was that this is a technicality dressed up as a principle. It is not, and the more I sat with it, the more coherent it looks.
Fair use has never been about whether you like the output. It asks whether the new use does something different from the original. A model that learns statistical relationships across millions of books is not competing with any one of those books the way a photocopied edition would. Courts looked at that and said yes, that is transformative.
Acquisition is a separate act with its own separate rules. If you break into a library at night and steal a book, writing a brilliant review of it afterward does not retroactively legalize the break-in. The review can be fair use. The burglary is still burglary. The courts drew that same line, and it holds up.
What it means in practice is that the legal exposure for AI companies is not really about the models. It is about procurement records. Where did the data come from. Who signed off. What is in the logs.
What this means if you build with these tools
I review AI tools for a living, and I read these rulings as a risk assessment document more than a philosophical one. A few things follow directly:
- Vendors that can document clean data sourcing have a real, defensible advantage. Not a marketing advantage. A legal one.
- Vendors that cannot document it are carrying a liability they may not have priced into your contract.
- “Trained on publicly available data” is not the same claim as “trained on legally acquired data,” and the gap between those two phrases is where the unresolved litigation lives.
- Indemnification clauses in enterprise AI agreements are about to matter a lot more than they did a year ago.
If you are picking a model provider for anything that touches your business, ask about data provenance. Ask in writing. The answer you get, or fail to get, tells you something about how that company handles risk generally.
Where the ambiguity actually sits
I want to be careful not to oversell these rulings. They are district court decisions. They are US decisions. Other jurisdictions are working through their own versions of the fair use question and are not obligated to agree. Cross-border discussions on fair use are ongoing precisely because there is no global answer here.
The piracy claims still going to trial could reshape the practical calculus even if the fair use holding survives untouched. A company can win the abstract legal principle and still lose badly on damages for how it built its dataset. Those are different outcomes with very different balance sheet consequences.
So when someone tells you the AI copyright question is settled, they are describing half of a two-part answer. The half about transformation looks stable. The half about where the files came from is still very much live.
My honest read
The framing I keep coming back to is that courts have decided AI training is a legitimate use of books and a completely separate question from whether you were allowed to have those books. That is not a loophole. It is two different laws doing two different jobs, and the industry spent years assuming the first one was the hard part.
It was not. The hard part is the receipts.
đź•’ Published: