The AI safety conversation has become unfalsifiable, and that’s a bigger problem than whatever the models are actually doing.
I review AI tools for a living. That means I spend most of my week separating marketing copy from function, and I’ve gotten reasonably good at it. But the safety discourse in 2026 has broken my usual method, because the usual method requires being able to check something. Right now, a huge share of what circulates as AI safety information can’t be checked by anyone outside a handful of labs.
TechCrunch put it plainly this month: two viral AI safety conversations demonstrated how hard it has become to tell AI fact from AI fiction. One involved Andrew Yang. That’s the level we’re operating at now, where the story isn’t a model doing something alarming, it’s people arguing about whether the alarming thing happened at all.
Unexpected behavior is doing a lot of work
Here’s what’s actually on the record. Safety discussions intensified in 2026 following unexpected model behaviors. ABC News and CBS News both ran segments on it in mid-September. Experts are warning about rapid advancement and potential threats.
Read that again and notice how little it tells you. “Unexpected behavior” is a category so wide it could mean a model produced a weird output in a research sandbox, or it could mean something genuinely worth losing sleep over. The phrase travels perfectly through news cycles precisely because it’s empty. Everyone can pour their preferred interpretation into it.
That vagueness isn’t accidental, and I don’t think it’s entirely malicious either. Labs have real reasons not to publish detailed accounts of model failures. But the effect is that the public conversation runs on summaries of summaries, and the loudest voices are the ones least constrained by having to be specific.
The talks that we know happened
One concrete thing did surface. On September 15, TechCrunch reported that OpenAI, Anthropic, and Google had been in talks about AI safety for weeks. Chris Lehane, OpenAI’s global policy chief, told reporters about it.
Take that seriously, because coordination between direct competitors is not free. These companies are fighting over the same engineers, the same enterprise contracts, the same developer mindshare. Sitting in rooms together on safety costs them something. That’s a signal worth more than a hundred thought-leader threads.
But notice who delivered the news. A policy chief, to reporters. Which means what we know about the most significant safety coordination in the industry arrives through the communications function of one of the participants. I’m not accusing anyone of lying. I’m pointing out that the only window we have into this process is one the participants built and control.
Why this wrecks tool reviews
My job is telling you whether an agent works, whether it’s worth your money, and where it breaks. Safety claims are supposed to be part of that assessment. In practice, they’ve become the least useful section of any review I write.
Ask a vendor about safety and you’ll get a page describing alignment work, red-teaming, and guardrails. None of it is verifiable from the outside. I can test whether an agent completes a task. So the safety section of a review becomes a summary of what the vendor said about itself, which is worth roughly nothing.
The gap gets filled by vibes. Some tools feel safe because the company has a serious reputation. Others feel reckless because the marketing is loud. Neither instinct is evidence, and I try not to dress either one up as analysis.
What I’d actually ask for
If labs want the safety conversation to mean something, the fix is specificity. Not more announcements about ongoing talks. Not more expert warnings about rapid advancement, which at this point function as background noise.
- Describe actual incidents with enough detail that outside researchers can evaluate them
- Publish what the joint discussions between OpenAI, Anthropic, and Google are producing, not just that they’re happening
- Give independent testers access to evaluate the claims that end up in marketing material
- Separate genuine safety research from positioning against competitors and regulators
Until some of that happens, my advice to anyone reading tool reviews, including mine, is to treat safety claims as unverified by default. Not false. Unverified. That’s a meaningful distinction and it’s the honest position.
The models may well be doing surprising things. Something clearly rattled these companies enough to get them talking. I just can’t tell you what, and neither can anyone else currently telling you they can.
🕒 Published: