Slow down, says the guy accelerating.
Dario Amodei has become convinced that the risks of AI need more prudence, and he’s written an essay saying so. His plan calls for industry-level restraint plus global regulation. Sam Altman has floated the idea that it may be time to “pace” AI development. Competitors have voiced support for third-party monitors who would evaluate model safety during development. On paper, that’s rare alignment among people who normally can’t agree on what a benchmark measures.
I review AI tools for a living. I test the things these labs ship, and I write down what actually happens rather than what the launch post promised. So when the people building the fastest models start asking for speed limits, my first question isn’t whether they’re sincere. It’s mechanical: who presses the pedal, and what happens to the person who presses it first?
The coordination problem nobody has solved
Voluntary pacing is a promise with no enforcement attached. If four labs agree to slow down and one doesn’t, the one that doesn’t wins the quarter. That’s not cynicism about anyone’s character, it’s just how competitive markets function when the prize is enormous and the penalty for restraint is losing.
Amodei knows this, which is presumably why his proposal reaches past industry agreement toward regulation, including at a global level. But that’s where the plan stops being a plan and starts being a wish. Global regulation of a software industry requires agreement between governments that can’t agree on trade, data privacy, or chip exports. The parties most likely to defect are also the parties least likely to sign.
The third-party monitor idea is the most concrete piece of this, and it’s the one I’d watch. Independent evaluation during development, rather than after launch, is a real mechanism. It’s also the piece with the most unanswered questions. Who staffs these monitors? Who pays them? What compute do they get access to? What authority do they have when they find something? An evaluator who can write a report but can’t stop a release is a compliance department, not a brake.
What “pacing” means in practice
Here’s what I keep running into as a reviewer. The gap between a model’s announced capability and its reliable, everyday behavior is already wide. Agents that promise autonomous multi-step work still get stuck, still hallucinate tool outputs, still need a human watching. That gap isn’t a safety feature. It’s a symptom of shipping fast and letting users absorb the difference.
So there’s a version of “pacing the frontier” that I’d welcome immediately, and it has nothing to do with existential risk. It’s slower releases, longer evaluation windows, honest capability documentation, and fewer demos that only work in the demo. That version is measurable. You can check it against what ships.
The version being discussed publicly is harder to verify. Prudence is a virtue that leaves no trace. A lab can say it is pacing while training at maximum capacity, and nobody outside can tell the difference without the kind of access that doesn’t currently exist.
The incentive question I can’t get past
When a frontier lab CEO calls for slowing down, there are two readings and both can be true at once. One: he genuinely believes the risk profile has changed and is saying so in public at some cost to himself. Two: calls for regulation from the leader in a space tend to shape rules that the leader can already meet, and that smaller competitors cannot.
I don’t think you have to pick. Sincere conviction and favorable positioning coexist comfortably. What I’d ask is that we stop treating the statement itself as the accomplishment. An essay is not a mechanism. Agreement in principle among CEOs is not oversight.
What would actually count as evidence
If pacing is real, it should be visible from the outside. A few things I’d take seriously:
- Third-party evaluators with contractual access to pre-release models and the power to delay a launch, not just publish concerns
- Published evaluation criteria that don’t change after the results come in
- A named, dated commitment from more than one lab, with a stated consequence for breaking it
- Capability documentation that describes failure modes with the same energy as the highlight reel
None of that requires a treaty. All of it is checkable. Until some of it exists, “pace the frontier” is a shared sentiment among the best-resourced people in the industry, which is a fine place to start and a terrible place to stop.
I’ll keep testing what ships. That’s the only pacing data any of us actually get.
🕒 Published: