Imagine a bakery that leaves its recipes in a binder on the counter. A guy walks in every morning, reads a few pages, walks out, and opens a shop down the street selling something that tastes suspiciously familiar. He never copies a recipe word for word. He just learned. And when the bakery asks for a cut, he explains that learning isn’t stealing, and besides, he’s already built forty thousand ovens.
That’s roughly where we are with the Seattle Times and Newsday, the two latest newspapers to sue OpenAI and Microsoft over copyrighted content allegedly used to train AI models without permission. They want damages. They also want their journalism pulled out of the training datasets. One of those requests is a normal legal ask. The other is closer to asking someone to un-bake a cake.
What they’re actually asking for
Attorneys for both papers argue OpenAI’s models are causing substantial harm to media businesses. That’s the core claim, and it’s not an unreasonable one to make in front of a judge. But I test AI tools for a living, and the part of this filing that interests me isn’t the damages figure. It’s the removal request.
Pulling specific content out of a trained model is not like deleting a row from a database. The text isn’t sitting in there as text. It’s been converted into weights, spread across billions of parameters, mixed with everything else the model ever read. There’s no folder labeled “Seattle Times” waiting to be dragged to the trash. The realistic options are retraining from a filtered dataset, which is expensive and slow, or bolting on output filters that try to stop the model from reproducing certain material, which is a patch rather than a fix.
So when a lawsuit asks for content removal, it’s really asking for one of two things: a very large bill, or a precedent that forces companies to license data upfront next time. Both are legitimate goals. Neither is technically what the words describe.
Why this keeps happening
These two papers are described as the latest to file, which tells you the pattern is now established. News organizations tried complaining, then tried negotiating, and a number of them have landed on litigation as the only lever that produces a response. That’s not activism. It’s a business decision made by people watching their traffic and their ad revenue behave badly.
From where I sit, testing chat assistants and agents week after week, the frustration makes sense even before you get to training data. The behavior that hurts publishers most is what happens at the output end. A user asks a question, the assistant produces a tidy summary sourced from reporting somebody paid for, and the user never clicks anything. The reporting still cost money. The reader still got what they needed. The transaction that used to fund the first part no longer connects to the second.
Training is the legal hook because copyright law has something to say about copying. The revenue problem is broader than that, and a court ruling on datasets won’t fully solve it.
What this means if you use these tools
Practically, in the short term, not much changes. Nobody’s assistant is going to stop working because a Long Island daily filed paperwork. But there are second-order effects worth watching if you build on top of these models:
- Licensing costs get passed along. Every deal signed to settle or avoid a suit becomes an input cost. That cost lands in API pricing eventually.
- Output gets more cautious. Models that fear regurgitation complaints tend to hedge, refuse, or paraphrase into mush. That’s a quality regression for anyone doing research or summarization work.
- Citation behavior improves. Legal pressure is the most reliable driver of source linking we have. Vendors don’t add attribution because it’s nice. They add it because it’s defensible.
- The data provenance question stops being optional. If you’re shipping something on top of a foundation model, “we don’t know what’s in it” is becoming an answer with liability attached.
My honest read
I’m not sentimental about newspapers, and I’m not impressed by AI companies acting surprised that the people they scraped are annoyed. Both sides made choices. The publishers left their work open to crawlers for two decades because search traffic paid the bills. The AI labs took that openness as consent and built products worth enormous sums without asking. Now everyone’s discovering that the old arrangement was a handshake, not a contract.
What I’d like to come out of this isn’t a giant payout. It’s a functioning market where using someone’s reporting to train a commercial model is a line item, not a gray area. That would be better for publishers, and honestly better for the tools I review, because models trained on properly licensed material can be documented, audited, and trusted in ways the current generation can’t.
Two more papers filed. Expect more. The interesting question isn’t whether these particular suits win. It’s whether anyone builds the licensing pipes before the courts start guessing at what they should have looked like.
🕒 Published: