Reflection AI’s valuation went from $8 billion to $25 billion in roughly six months. That’s a 3x markup on a company whose flagship open-weight model, as of early March 2026, had not shipped publicly. Now Axios reports a powerful new model is coming. No date. No benchmarks. No details.
I review AI tools for a living. My job is to download the thing, break it, and tell you whether it’s worth your afternoon. So I want to be upfront about my bias here: I have nothing to review. And that’s the entire story.
What’s actually confirmed
Let’s separate the money from the product, because they’re moving at wildly different speeds.
- October 2025: $2 billion raised at an $8 billion valuation.
- March 2026: another $2 billion, valuation jumps to $20 billion.
- April 2026: CEO Misha Laskin confirms the latest round closed at a $25 billion pre-money valuation.
- June 2026: a multiyear compute agreement with SpaceX potentially worth up to $6.3 billion.
- July 2026: Bloomberg reports Nebius will sell $1 billion in AI capacity to Reflection.
- Backed by Nvidia. Positioned, per the New York Times, to compete with DeepSeek.
On the product side: the frontier open-weight model at the center of the pitch was still unreleased as of early March 2026, and the code research agent Asimov sat behind a waitlist. Axios now says a powerful new model is set to arrive. Release date undisclosed.
Compute commitments are not a product
Here’s what the SpaceX and Nebius deals tell me, and what they don’t. They tell me Reflection has secured serious capacity and has the balance sheet to pay for it. Up to $6.3 billion in one agreement plus a billion in another is not a rounding error, and compute providers don’t sign multiyear paper with companies they expect to evaporate.
What those deals don’t tell me is whether the model is any good. Buying GPUs is a purchasing decision. Shipping a frontier open-weight model that developers actually prefer over the alternatives is an engineering outcome. The industry keeps conflating the two, and investors keep rewarding the conflation.
I’ve tested enough tools to know the pattern. Enormous compute spend correlates with capability, sure. It does not guarantee it. And it says nothing about the parts that determine whether you’ll use the thing: latency, tool-calling reliability, how it behaves on long context, whether the license is actually permissive or just marketed that way.
The open-weight promise is the real test
Reflection’s whole pitch rests on open weights. That’s a meaningfully different bet than the closed frontier labs are making, and if it lands, it matters more than another incremental API release. Open weights mean you can run it, fine-tune it, audit it, and deploy it without asking permission or routing your data through someone else’s servers.
It also means there’s nowhere to hide. A closed model can be tuned, patched, and quietly adjusted after launch. Weights you’ve downloaded are frozen evidence. Every community benchmark, every red-team probe, every “actually this falls apart on multi-step reasoning” thread hits all at once, publicly, within days.
That’s a brutal review environment, and it’s the correct one. The DeepSeek comparison in the NYT coverage is apt, and it cuts both ways. DeepSeek earned its reputation by shipping weights that held up when strangers stress-tested them. Reputation in open models is downloaded, not announced.
What I’ll be looking for
When this model actually lands, here’s the checklist I’ll run, and I’d suggest you hold it to the same standard:
- Are the weights genuinely downloadable, or is it another waitlist with open-source branding?
- What’s the license? “Open” covers everything from Apache 2.0 to restrictions that make commercial use a legal problem.
- Does Asimov come off the waitlist, and does the agent work on real repositories instead of curated demos?
- How does it hold up on independent evaluations rather than first-party charts?
- What does it cost to actually run at the parameter count they ship?
My honest read
Reflection has bought itself an extraordinary runway and an equally extraordinary expectations problem. At $25 billion with Nvidia’s backing and billions in compute lined up, a merely good model reads as a disappointment. The funding has already priced in a win.
I’m genuinely interested. A well-executed open-weight frontier model from a well-capitalized US startup would change what developers can build without renting intelligence by the token. But interest isn’t endorsement, and an Axios scoop about an unreleased model with undisclosed specs is a press cycle, not a product.
Ship the weights. I’ll have a review up the same week.
🕒 Published: