\n\n\n\n Somebody Else's Breakthrough Was My Tuesday Last Year - AgntHQ \n

Somebody Else’s Breakthrough Was My Tuesday Last Year

📖 4 min read•793 words•Updated Sep 19, 2026

In September 2026, a frontier lab called TypeSafe AI — founded by Diogo Almeida, a co-inventor of ChatGPT at OpenAI — launched a system named Jev and framed the idea behind it as a breakthrough. The idea: non-autoregressive decision models trained with reinforcement learning. Decisions produced in one shot instead of token by token.

I built that a year earlier. It’s called Laya. You can pip install laya right now. There’s a GitHub repo and a live demo. It runs a multilingual System 1 decision engine at 33 milliseconds with calibrated probabilities.

So let me be honest about how that feels, and then let me be honest about why it doesn’t actually matter as much as my ego would like.

Priority is a story people tell after the fact

Every engineer who has shipped something early has lived this moment. You solve a problem, you publish it, you get a handful of stars and a couple of GitHub issues from strangers, and then eighteen months later the same architecture shows up with a launch video and a founding team whose résumés open doors yours doesn’t.

The temptation is to yell about it. I’ve watched enough of these fights on this site to know how they end: nobody’s mind gets changed, and the person with the bigger distribution wins the narrative by default.

What I’d rather do is point at the thing both of us apparently figured out independently, because convergence is the actual signal here. When two unrelated efforts land on the same architecture, that architecture is probably correct. Independent discovery is stronger evidence than either project’s marketing.

Why one-shot decisions are the right bet

Autoregressive generation is a beautiful hammer and we’ve been swinging it at problems that aren’t nails. If a model’s job is to write an essay, generating one token at a time makes sense — each word conditions the next. But if the job is to answer a question with a bounded set of outcomes, sequential generation is overhead you’re paying for nothing.

Routing a support ticket. Deciding whether a tool call is warranted. Classifying intent before you hand off to an expensive model. Flagging whether a piece of text is safe. These are decisions, not compositions. You don’t need the model to think out loud in order to produce a label and a confidence score.

This is where the latency number matters. 33 milliseconds is not “fast for an LLM.” It’s fast in the sense that it disappears inside the rest of your stack. You can put a decision like that in a hot path and not think about it again. Anything that requires a round trip to a generative model, plus a chain of thought, plus parsing, does not belong in a hot path, no matter how good the benchmarks look.

The calibration part nobody wants to talk about

If I had to pick one thing that separates a useful decision model from a demo, it’s calibrated probabilities. Not a label. Not a vibe. An actual number that means something when you compare it to the next number.

Agent frameworks are full of components that return confident answers with no usable confidence. That’s how you end up with routing logic built on string matching and prayer. If your classifier says 0.6 and that genuinely means it’s right about six times out of ten, you can build real control flow on top of it:

  • Route high-confidence cases to the cheap path and let the expensive model handle the rest
  • Set escalation thresholds that reflect your actual tolerance for being wrong
  • Track drift by watching whether calibration holds, not just whether accuracy holds
  • Give humans a meaningful queue instead of a random sample of everything

That’s the part I’d tell people to evaluate first when a new decision engine shows up, mine included. Speed is easy to measure and easy to fake with a small model. Calibration is harder to fake and more useful to have.

What I’d actually ask you to do

Don’t take my word that I got there first. Take the architecture seriously enough to test it against whatever you’re currently using for classification and routing. Run Laya. Run Jev. Run both against your own data, because your own data is the only benchmark that has ever mattered.

If the frontier lab’s version is better, use theirs. I’d rather the idea win than the credit. The thing I care about is that agent builders stop reaching for a generative model every time they need a yes-or-no answer. That habit is expensive, slow, and produces systems nobody can reason about.

Being early is not a competitive advantage. Being early and reachable is. Laya has been installable and documented since before this idea became a launch announcement. That’s the version of vindication I’ll accept.

🕒 Published:

📊
Written by Jake Chen

AI technology analyst covering agent platforms since 2021. Tested 40+ agent frameworks. Regular contributor to AI industry publications.

Learn more →
Browse Topics: Advanced AI Agents | Advanced Techniques | AI Agent Basics | AI Agent Tools | AI Agent Tutorials
Scroll to Top