The most interesting thing about Kev is not that it copies Jev. It’s that you can actually check it yourself, which puts it ahead of roughly everything else written about this whole category in the past week.
Let me set the scene. Jev, a “decision model” that trades free-form text generation for something more constrained, went viral in September 2026. Within 48 hours, at least six Jev clones shipped. Coverage followed immediately — including, by its own admission, coverage at explainx.ai that was assembled from digest headlines with no primary source to check against. That’s not a knock on one site. That’s the shape of the entire news cycle around this thing: a lot of confident writing about models nobody had run.
Kev is Jared Palmer’s entry. It’s a small family of decision models built on Qwen3.5, similar in spirit to Jev but using different training methods. Three sizes: 0.8B, 4B, and 9B parameters. Released in 2026. You can train it and run it on your own hardware. It hit 370 points and 164 comments on Hacker News inside 14 hours.
That’s the fact sheet. Everything else you’ve probably read about it is inference, mine included. So let me be clear about which is which.
Why the size range is the actual story
0.8B, 4B, 9B. Not one flagship checkpoint with a press release. A ladder.
That choice tells you what the project is for. A 0.8B model is not competing with anything at the frontier — it’s small enough to run on a laptop, small enough to fine-tune without renting a cluster, and small enough that you can retrain it repeatedly while you figure out whether the idea even works for your use case. The 9B sits at the other end of what a single consumer GPU can reasonably hold.
A release shaped like this is aimed at people who want to reproduce and modify, not people who want to call an API. That’s a different audience than most model launches court, and it’s a more honest one. Reproducibility is the claim that’s easiest to falsify, which is exactly why it’s the claim worth making.
What I can’t tell you
I’m not going to give you performance numbers, because I don’t have verified ones. The repo carries planning documents — PLAN.md, PLAN_Qwen35.md — with dated entries into September 2026 and fragments referencing MMLU-Pro figures. Working notes are not results. A number sitting in a living plan file could be a target, a baseline, a reference point for a different model entirely, or a stale line someone forgot to delete. Treating it as a benchmark score would be exactly the mistake the rest of this news cycle already made.
Here’s what I’d want before recommending Kev to anyone:
- Published evaluations on the decision tasks it’s actually meant for, not general knowledge benchmarks borrowed from chat models
- A clear description of how its training method differs from Jev’s, and what that difference buys you
- Reproduction reports from people who are not the author
- Some sense of what happens at 0.8B versus 9B, because the gap between those two is enormous and the use cases are not interchangeable
None of that is a criticism. It’s a to-do list for anyone evaluating this seriously, and it’s the work that gets skipped when six clones ship in two days.
The clone wave deserves more skepticism than it’s getting
When a viral model spawns half a dozen reimplementations inside 48 hours, the natural read is momentum. My read is different. Fast clones are cheap to produce and expensive to validate. A repo, a config, and a base model get you to “announced” very quickly. Getting to “this reliably does the thing” takes considerably longer, and the news cycle has already moved on by the time anyone finds out.
Kev’s advantage in that crowd is not that it’s better. I have no evidence for that. Its advantage is that it’s open, small enough to verify, and paired with visible planning docs. You can disprove it. In a week where the secondary coverage was reportedly built on headlines rather than sources, being checkable is a real differentiator.
Where I land
If you’re building something that needs constrained, local decision-making and you have the appetite to train and evaluate a model yourself, Kev is worth a weekend. Start at 0.8B, write your own eval set for your actual task, and decide from your own numbers.
If you’re looking for a drop-in replacement with proven behavior, wait. Not because Kev looks weak, but because nobody has published the evidence either way, and the loud consensus forming around this category is running well ahead of what anyone has actually measured. That gap tends to close in one direction or the other fairly quickly. Let someone else find out first.
đź•’ Published: