The moat moved.
Not the weights. Not the architecture. Not even the raw compute, though plenty of people are still writing checks as if it were. In 2026, the thing that actually separates a useful model from an expensive one is what happens after pretraining: the post-training data, the reward model that scores it, and the infrastructure that runs the reinforcement learning loop without setting money on fire.
That’s the context for the current wave of interest in open-weight decision models and the RL fine-tuning platforms springing up around them. I’ll be straight with you about Clef specifically: the verified details I have in front of me are thin. I’m not going to invent benchmark numbers or quote a founder who may not have said anything. What I can do is tell you what the category looks like right now, and what questions you should be asking before you hand anyone your data.
Why open weights stopped being a consolation prize
For a couple of years, “we use open models” was code for “we couldn’t afford the good ones.” That framing is dead. The open-weight path now holds up as a real alternative in specific domains, and the common thread in those domains is a proprietary reward signal. If you have a way to score outputs that nobody outside your building can replicate, you have something the frontier labs can’t buy.
DeepSeek’s GRPO work is a big part of why this is tractable. Group relative policy optimization made RL on open weights something a competent team could actually run, rather than a research project with a nine-figure budget. The technique spread fast, and the tooling followed.
The clearest proof point I’ve seen is Bridgewater Associates, which used Tinker, a cloud fine-tuning service, to adapt open-weight models to its own data. A hedge fund with decades of proprietary signal doesn’t need a general-purpose chatbot. It needs a model that’s excellent at its particular decisions. That’s the whole argument in one example.
Post-training is where the gains actually live
On the closed side, the same story plays out. Anthropic’s Opus 4.7 posted improvements on SWE-Bench Verified and Pro, and the lever was post-training work pairing constitutional AI with RL. Same base-model game, different finishing process, measurably better results.
Take the two together and you get the 2026 picture: whether your weights are open or closed, the differentiation comes from three things.
- Post-training data you own and nobody else can scrape
- Domain-specific reward models that encode what “good” means in your business
- RL compute infrastructure that makes iteration cheap enough to run many times
Everything else is table stakes. Meanwhile the capital keeps flowing toward raw scale, with more $60 billion-class compute partnerships and infrastructure deals looking likely. Both bets can pay off. They’re just different bets.
The platform pitch, and where it gets slippery
Fireworks is selling reinforcement fine-tuning on the promise that you can train expert open models that surpass closed frontier models, with two weeks of free training to get you started. The free trial is a smart offer and a good signal of confidence. It’s also the part where I’d slow down.
Kyle Corbitt of OpenPipe has been loud about the unglamorous parts of RL fine-tuning, and his list of hazards is the one that matters: rubrics, environments, and reward hacking. Those three words explain most failed RL projects. Your model doesn’t learn what you meant. It learns what you scored. If your rubric has a loophole, the model will find it faster than your reviewers will, and the resulting metrics will look terrific right up until production.
So before you sign anything, ask:
- Who writes the reward model, you or the vendor? If it’s the vendor, you’ve outsourced your moat.
- Can you inspect and version the environment the model trains in?
- What happens to your training data and the resulting weights if you leave?
- How many full training runs can you afford per month once the free window closes?
Small, specialized, and cheap is winning
The economics reinforce all of this. Open models keep closing the quality gap while costing dramatically less to run, and the releases keep coming: Kimi K2.7 Code shipping day-zero on Fireworks for agent work at lower cost per task, MiniMax M3 offering long context and native multimodality at a twentieth of the price. Simon Willison’s collaborator Laurie Voss has called 2026 the year of fine-tuned small models, and the pricing curve backs him up.
My read for anyone evaluating an open-weight decision model or a new RL platform: the technology is ready, the tooling is good enough, and the hard part is entirely yours. If you can’t articulate what a correct decision looks like with enough precision to score it automatically, no platform fixes that. If you can, you may be sitting on the only durable advantage left.
🕒 Published: