It’s 2:40 in the morning and you’re staring at a JSON blob that a language model swore was valid. It isn’t. There’s a trailing comma where a boolean should be, the confidence score came back as the string “high” instead of a float, and somewhere in the middle of that response the model decided to explain its reasoning in three friendly paragraphs you never asked for. You wrote a regex to strip them out last week. It broke tonight.
If that scene feels familiar, you’re the exact person Jev was built for.
On September 15, 2026, a company called TypeSafe AI announced a frontier model that cannot write a sentence. Not “prefers not to.” Cannot. Jev returns typed probabilistic decisions for software instead of generating text. The person behind it is Diogo Almeida, a former OpenAI researcher who helped build ChatGPT and co-invented reinforcement learning from human feedback, arguably the technique that made chatbots tolerable to talk to. He spent roughly two years working on this quietly. His pitch is that today’s large language models are structurally inefficient, and that Jev is faster and more efficient as a result.
I want to be careful here, because this is the part of the news cycle where everyone loses their mind. So let me separate what we actually know from what’s being projected onto it.
What’s genuinely interesting
The framing is the smart part. For three years, the industry’s answer to “I need structured output” has been to take a machine built for open-ended prose and squeeze it through a funnel. Function calling, JSON mode, grammar-constrained decoding, retry loops, validation layers, schema coercion. Every one of those is a patch on a mismatch. You wanted a decision. You got an essay, and then you paid tokens to have the essay pretend to be a decision.
Almeida’s claim that LLMs are structurally inefficient lands differently coming from him than it would from a random startup founder. He built the thing. He built the training method that made the thing feel human. If the guy who taught models to chat is now saying the chatting is the overhead, that’s at minimum worth reading twice.
The secondary claim floating around the coverage is that a model like this can’t hallucinate. I’d treat that phrasing with real suspicion, and I’ll explain why in a second, but the underlying logic isn’t nonsense. A model that outputs a typed value inside a defined space has a much smaller surface on which to invent things. It can’t cite a fake case law. It can’t apologize and then repeat the same error. It can’t drift into a tone. The failure modes shrink because the output space shrinks.
What I’m not buying yet
“Can’t hallucinate” is a category trick. A model that returns a probability distribution over valid options can still be confidently, catastrophically wrong. It just gets to be wrong in a well-formed way. A 0.94 confidence score attached to the wrong branch is not a hallucination in the token sense, but it will absolutely ruin your afternoon. Type safety guarantees the shape of an answer, never the truth of it. Any developer who’s shipped a strongly typed system that returned garbage already knows this.
Second, “faster and more efficient than previous models” is a claim, not a benchmark. I haven’t seen independent numbers. Efficiency against what baseline, at what task, with what accuracy floor? Comparing a decision model to a general text model on speed is a bit like comparing a calculator to a laptop. One of them is faster because it does less. That can be exactly the right tradeoff, but it isn’t the same as a breakthrough in capability.
Third, and this is the one nobody wants to say out loud: narrow models have a distribution problem. Developers adopt what’s already in their stack. A model that only does one thing has to be dramatically better at that one thing to justify a new vendor, a new SDK, and a new line item. “Slightly cleaner than JSON mode” won’t clear that bar. “Ten times cheaper per decision with measurable reliability gains” would.
Why the developer excitement makes sense anyway
Agent builders have been the most quietly miserable group in AI for two years. Their systems fail at the joints, where one model’s fuzzy text output becomes another component’s structured input. Every handoff is a coin flip dressed up as a pipeline. A model designed for that specific job, from someone with this specific résumé, is the first credible attempt at fixing the joints instead of adding another retry.
That’s the real story. Not a new frontier model. A bet that the next useful thing in AI is smaller and quieter than the last one.
I’ll withhold judgment until someone outside TypeSafe AI publishes numbers. But I’d rather review a model that admits what it doesn’t do than one more chatbot that claims it does everything.
🕒 Published: