Every car company builds a halo model. The one with the carbon fiber body panels and the engine that costs more than a house. It exists to sit in the showroom, win awards, and make you feel good about buying the sedan. Almost nobody actually buys the halo car. That was never the point.
Anthropic appears to have accidentally built one in software. According to reporting from the Financial Times, picked up and amplified by Gary Marcus, the company’s best model is struggling to attract users while cheaper tools thrive. Overall growth is strong. Adoption of the flagship specifically? Low, compared to the budget options sitting right next to it on the menu.
I’ve been reviewing AI tools long enough to find this completely unsurprising and still slightly delicious.
The assumption that just broke
For roughly three years, the industry has operated on a simple article of faith: build the smartest model, and demand follows. Capability is the moat. Benchmarks are the marketing. Whoever sits at the top of the leaderboard collects the revenue.
This story punches a hole in that. It suggests something the frontier labs really do not want in a pitch deck, which is that most people’s actual work does not require the most capable model available. It requires a model that is good enough, fast enough, and cheap enough to run a thousand times a day without anyone checking the bill.
Those are different products. They’ve been priced as if they’re the same product with a quality dial.
What I see in actual tool reviews
When I test agents and AI-powered tools, the pattern repeats constantly. A team builds something clever, wires it to the best available model, demos beautifully, then quietly downgrades before launch because the unit economics don’t survive contact with real usage. The output difference is often noticeable in a side-by-side comparison and invisible in a production workflow where the model is summarizing a support ticket or reformatting JSON.
The flagship wins the eval. The cheap model wins the deployment. This has been true for a while in the tools I look at, and it now appears to be showing up in the numbers at the model layer.
A few things I’d watch for if you’re making buying decisions:
- Tier your tasks, not your vendor. Route the hard reasoning to the expensive model and everything else to the cheap one. Teams that do this cut costs dramatically without users noticing.
- Test on your work, not benchmarks. The gap between models on a public eval tells you almost nothing about the gap on your specific, boring, high-volume task.
- Price the ceiling, not the average. Cheap models let you retry, verify, and run multiple passes. Sometimes three cheap attempts beat one expensive one.
Why this matters beyond Anthropic
Marcus frames this in the context of upcoming IPOs, and that framing is the sharp part. The valuations attached to frontier labs assume that frontier capability converts to frontier pricing power. If the market’s appetite sits mostly in the cheap tier, that’s not a small adjustment to the model. It’s a different business with different margins and different competitors, many of whom are giving away perfectly usable models for free.
There’s an awkward wrinkle here too. Anthropic has publicly said that as of May 2026, more than 80% of the code merged into its own codebase was authored by Claude. The internal case for the best model is real and demonstrated. Translating that into external demand at premium prices is a separate problem, and this reporting suggests it’s an unsolved one.
My read
I don’t think this means the frontier is pointless. Somebody has to build the halo car, and the research that produces it trickles down into the cheap models everyone actually uses. That’s how the automotive analogy resolves, and it’s probably how this one does too. Today’s flagship becomes next year’s default at a fraction of the cost.
But it does mean the pitch needs to change. “We have the best model” is not a business strategy if the best model is a showroom piece. The interesting question is whether the labs can make premium capability genuinely necessary for enough real work to justify what they’re charging, or whether they end up as suppliers of intelligence that gets commoditized within eighteen months of release.
For anyone buying AI tools right now, the practical takeaway is simpler and slightly liberating. You are almost certainly overpaying for capability you cannot detect in your own output. Go run the comparison. The cheap tier has gotten embarrassingly good, and the market seems to have figured that out faster than the marketing did.
🕒 Published: