What happens to your favorite AI coding assistant when a bad output kills someone?
That question doesn’t come up much in the tool reviews I write. Most agents I test fail in boring ways. A wrong import. A hallucinated API method. A confidently broken migration script that takes down a staging database nobody was using anyway. You shrug, you re-prompt, you move on. The entire consumer AI economy runs on the quiet assumption that mistakes are cheap and retries are free.
Three companies are about to sit on a stage at TechCrunch Disrupt 2026 and talk about what happens when that assumption doesn’t hold. Shield AI CTO Nathan Michael, Waabi founder and CEO Raquel Urtasun, and General Motors Director of Robotics Strategy Mikell Taylor are taking the Real World AI Stage to discuss safety validation, regulatory navigation, and trust-building in high-stakes deployment. Defense autonomy, self-driving trucks, and production vehicles. Three domains where a failed output has a body count.
The stat that should make every engineering manager uncomfortable
GM disclosed that nearly 90% of the code created by its autonomous driving team is now AI-generated.
Read that again with your reviewer brain on. Not 90% of internal tooling. Not 90% of test scaffolding. The code from the team building systems that steer two-ton machines past pedestrians. And the number is high enough that “we use AI to help with boilerplate” no longer covers it. That’s the primary authoring method.
My first reaction was skepticism, and I’ll keep some of it. “AI-generated” is a slippery category. Tab-completion counts. A function body scaffolded from a docstring counts. There’s a wide gap between an agent proposing a diff that three engineers then tear apart, and an agent committing to main. Nobody should read that number as “GM fired its programmers.”
But my second reaction is the interesting one. If you accept the figure at face value, the meaningful work has already moved. It didn’t move to prompting. It moved to verification.
Validation is the actual product
Here’s what I think this panel is really about, and why it matters more to agent buyers than the average Disrupt session.
In the consumer agent market, evaluation is an afterthought. Vendors ship benchmark scores that were cherry-picked in a lab, a handful of demo videos, and a changelog. When I ask how a tool behaves on a codebase it wasn’t trained to impress, I usually get marketing. There is no equivalent of a crash test. There’s barely an equivalent of a unit test.
Safety-critical teams don’t have that luxury. Regulators show up. Insurers show up. Simulation infrastructure isn’t a nice-to-have, it’s the thing you’re actually selling. Waabi has built its identity on that premise, and the market has responded: the company closed $1 billion in new funding in January 2026 and announced crossing what it calls the next frontier of generalization in June. You don’t raise that kind of money on a model checkpoint. You raise it on a claim that you can prove your system works before it touches a public road.
Shield AI operates under a similar constraint with different stakes. Defense autonomy means adversarial conditions, degraded communications, and consequences that don’t get patched in the next release.
Three companies, three markets, one shared bottleneck: how do you know the thing is safe when the thing was partly written by another thing you also can’t fully explain?
What this should change about how you evaluate agents
I don’t think most readers of this site are shipping autonomous trucks. But the transfer is real, and it’s the reason I’ll be paying attention to what these three actually say versus what reads well in a press summary.
- Ask about the verification layer, not the model. Any vendor can tell you which frontier model they wrap. Fewer can tell you what happens to a generated diff before it lands.
- Treat “percentage of code AI-generated” as a neutral metric. It measures authoring volume, not quality. High numbers are only impressive alongside evidence of how the output gets checked.
- Watch who invests in simulation. In safety-critical AI, testing infrastructure is the moat. In agent tooling, it’s usually the missing piece.
- Regulatory pressure is a feature. Domains with real oversight develop real measurement. The unregulated ones develop better demo reels.
The part I’ll be watching for
Panels like this tend to produce two kinds of moments. The first is a confident story about process maturity that nobody can independently check. The second, rarer and far more useful, is someone admitting which failure mode still keeps them up at night.
I want the second one. Specifically: when 90% of your code comes from a generator, what does code review even mean? Who owns a defect that no human typed? Does your validation suite catch a class of bug that human authors never would have produced in the first place?
Those questions don’t have clean answers yet. The companies being forced to answer them first are the ones where failure isn’t a retry. Everyone else in this industry is running the same experiment with the consequences turned off, and calling the silence success.
🕒 Published: