Zero. That’s how many new chips Kog is building to attack the AI inference cost problem. In an industry where the default answer to every performance question is “buy more accelerators,” a French startup is making the opposite bet: the hardware you already own is being wasted, and better software can fix it.
I’ll be honest — my inbox is full of pitches from companies claiming they’ve cracked inference efficiency. Most of them are wrappers around someone else’s work with a pricing page stapled on. Kog caught my attention for a different reason: they’re not selling a new chip, a new cloud, or a new model. They’re going deeper into GPU software itself to squeeze more inference out of silicon that’s already deployed.
The Contrarian Bet
The most interesting part of Kog’s pitch isn’t the efficiency angle — it’s the heresy. There’s a widely accepted belief in AI circles that GPUs are poorly suited for agentic workflows, those multi-step processes where an AI system reasons, calls tools, and chains actions together. Kog says the industry has this wrong. That’s not a small claim. It’s the kind of statement that either ages beautifully or becomes a cautionary slide in someone’s conference deck.
If Kog is right, the implications are worth taking seriously. Agentic workloads are where the industry is heading, and if existing GPUs can handle them better than assumed, a lot of infrastructure spending math changes. Companies wouldn’t need to wait for specialized hardware or pay a premium for exotic accelerators. They’d need better software sitting between their models and their chips — which is exactly what Kog is building.
Why Software Over Silicon Makes Sense
Let me give credit where it’s due: optimizing existing hardware is the unsexy, correct answer to a real problem. GPU costs keep climbing, and the standard response — throw more accelerators at it — is a treadmill. Every efficiency gain from software is a gain you don’t pay for twice. New chips take years to design, fabricate, and ship. Software improvements ship on your existing fleet.
This is also a more honest business model than most of what I review. Kog isn’t asking you to rip out your stack or migrate to a proprietary platform. The pitch is straightforward: same hardware, more inference, lower costs. Whether they can deliver at scale is the open question — and I don’t have benchmark numbers to evaluate yet, so I’m not going to pretend I do. But the approach itself is sound.
The European Angle
There’s a strategic layer here that shouldn’t be ignored. Kog plans to feed its methodology into agent-based pipelines that will let it support more chips and models over time. That expansion path matters, because Europe is actively trying to build its own capability in both chips and models. A French company that makes existing GPUs work harder fits neatly into that ambition — it’s efficiency as sovereignty, not just efficiency as cost savings.
Europe has struggled to produce AI infrastructure players that matter globally. Most of the serious GPU software work happens inside American chip giants or American AI labs. If Kog can carve out a real position in the inference optimization layer, it becomes strategically relevant beyond its own revenue line.
My Honest Take
Here’s where I land. The thesis is solid: software optimization on existing hardware is the highest-return, lowest-risk path to cheaper inference, and challenging the GPU-agentic-workflow assumption is exactly the kind of contrarian technical bet that occasionally pays off enormously.
What I want to see before I get excited:
- Independent benchmarks. Show me third-party numbers on real agentic workloads, not curated demos.
- Breadth of support. The roadmap toward more chips and models is the right one — execution on it is what separates a product from a research project.
- Proof at scale. Squeezing efficiency out of one GPU in a lab is different from doing it across a production fleet.
The inference cost problem isn’t going away, and the “just buy more GPUs” era can’t last forever — the economics don’t allow it. Someone will win the software optimization layer. Kog is making an early, foc
🕒 Published: