Zero. That’s how many world model companies have handed me enough detail to actually evaluate what they’re building. Not a partial spec, not a limited benchmark, not a “here’s what our model can’t do yet.” Zero. I review AI tools for a living, and this is the first category I’ve covered where the pitch decks are louder than the products and the technical disclosure rounds down to nothing.
The reporting on this has converged on the same read: world model companies are keeping a lot of secrets. Part of why is structural. As TechCrunch put it, the mystery comes from how versatile world models are as an idea. The simplest version is a navigable map of the world, similar to the AI models that power self-driving cars. But the same underlying concept stretches across robotics, AI software, and whatever else a founder decides to gesture at during a fundraise.
Versatility is a convenient fog
Here’s what bothers me as a reviewer. When a category is defined loosely enough to mean “a model that understands how physical reality behaves,” every company in it gets to claim progress without ever agreeing on what progress looks like. A navigable spatial map is a testable thing. You can measure whether it predicts where a car should be in two seconds. A general-purpose simulator of reality is not testable in any way you or I can check from the outside.
That gap is doing a lot of work right now. Vagueness is genuinely useful when you’re building something hard and early. It’s also genuinely useful when you’re building something less impressive than your valuation implies. From the outside, those two situations look identical, and that’s the whole problem.
The excuses, ranked by how much I buy them
Companies in this space have real reasons for staying quiet. Some hold up better than others:
- Global competition. The most credible one. AI has become a contest between major economies, with governments treating model capability as strategic. If your competitor is a national program rather than a startup down the road, you don’t publish your architecture. Fair.
- Cybersecurity. Also reasonable. Detailed disclosure about training pipelines and weights is a map for anyone looking to steal or attack them.
- Containment difficulty. Leading AI companies are openly struggling to control the behavior of their latest models. If you can’t reliably predict what your system does, saying less publicly is the safer move. Understandable, though not exactly reassuring.
- “It’s too early to share.” The weakest one. Early-stage teams have published rough, embarrassing, honest work for decades. Choosing not to is a decision, not a constraint.
Notice that the first three explain why companies withhold implementation details. None of them explain why so few will state plainly what their model does today, where it fails, and what it can’t handle. Those are different categories of information, and the space keeps treating them as one.
What this costs the people actually buying
If you’re a robotics team or a developer deciding whether to build on someone’s world model, secrecy shifts all the risk onto you. You can’t compare two vendors on any shared measure. You can’t tell whether a failure in your pilot is your integration or their model. You can’t estimate whether the thing improves on a schedule that matters to your roadmap. You’re buying a narrative and hoping the engineering underneath matches.
The honest workaround is unglamorous. Insist on running your own evaluation on your own data before signing anything. Ask specifically what the model gets wrong and treat a non-answer as an answer. Ask what happens when conditions fall outside the training distribution. Write your contract assuming capability claims are aspirational, because until someone shows their work, they are.
My actual position
I don’t think world models are hype. The idea that useful AI needs some internal model of how things move and interact is reasonable, and the self-driving lineage gives it a real technical foundation. What I’m skeptical of is the current information asymmetry, where an entire category has agreed that nobody has to prove anything yet.
Secrecy is defensible for weights, training data sources, and hardware specifics. It is much harder to defend for capability boundaries and failure modes, which is exactly the information customers need and exactly the information nobody is offering.
Until that changes, my review of this category stays the same: promising, unverifiable, and priced as if it were the first without being the second. I’d love to update it. Someone just has to show me something I can test.
🕒 Published: