$312 million. That’s what OLIX pulled in during its 2026 Series B, and I’d bet most people reading AI chip coverage this year couldn’t tell you what the company makes without looking it up. Meanwhile, Cerebras’ CS-4 claimed 30x faster inference than GPUs and Google’s Ironwood chip beat Nvidia’s Blackwell on performance, and those two stories ate the entire news cycle.
I review AI tools and agents for a living, which means I spend a lot of time watching people get excited about the wrong layer of the stack. The pattern in silicon is identical to the pattern in software: everyone argues about the model, nobody asks about the plumbing. And in 2026, the plumbing has a name that shows up in trade press and almost nowhere else — ABF substrate.
The part of the accelerator that isn’t the chip
Here’s what I mean by “the other chip.” Every AI accelerator you’ve read a benchmark about is not a single piece of silicon. It’s a die, or a stack of dies, mounted on a substrate — the layered organic material that routes thousands of electrical connections between the processor and the board it lives on. Tom’s Hardware ran a piece this year on the state of ABF substrates in data center silicon. That headline got a fraction of the attention of the Ironwood-versus-Blackwell story. Same supply chain. Same physical products.
I’m not going to pretend I have exclusive numbers on substrate capacity, because I don’t, and I’m not in the business of making them up. What I can tell you is what the funding pattern implies. Money moving into the unglamorous layers of a stack is usually a signal that the glamorous layers are constrained by them. Investors don’t put nine figures into a component category because it’s exciting. They do it because someone showed them a chart where demand outruns supply.
Why this matters if you never touch hardware
You might be thinking this is a semiconductor story and you build agents, so who cares. Fair. Let me connect it.
Every pricing decision in the AI tools you use traces back to compute availability. When inference gets cheaper, the tools you’re paying $200 a month for start looking overpriced, and the honest ones cut their rates. When compute gets tight, everything you depend on quietly degrades — rate limits tighten, context windows shrink on the cheaper tiers, latency creeps up during business hours. You’ve felt this. You probably blamed the vendor.
A claim like 30x faster inference only reaches you as a price cut or a latency improvement if the parts that make those accelerators physically exist can be built at volume. A performance win on a spec sheet and a performance win in your API bill are two different events, sometimes separated by a year or more of packaging and assembly capacity catching up.
The benchmark theater problem
This is the part where I get annoyed. The AI chip news cycle in 2026 is structured almost entirely around comparative performance claims, and comparative performance claims are the easiest thing in the world to stage. Pick the workload. Pick the precision. Pick the competitor’s configuration. Publish the multiple.
I’m not saying the Cerebras or Google numbers are wrong. I have no evidence they are, and I’d rather say “I don’t know” than perform skepticism I can’t back. What I’m saying is that a number like 30x tells you nothing about whether you’ll ever be able to rent that hardware, at what price, or in what quantity. Those questions get answered in substrate factories and advanced packaging lines, and almost nobody covers those.
Applied to tools, the same rule holds. A demo video is a benchmark. A funding round is not a product. The interesting question about any AI company is never “how fast is the best case” — it’s “what breaks first when you scale it.”
What I’d actually watch
If you want a useful signal from the chip space without pretending to be an analyst, watch three things:
- Where the money goes when the headlines aren’t looking. OLIX raising $312 million in a year dominated by Google and Nvidia news is more informative than either of those stories, because it tells you what people with capital think is scarce.
- Whether performance claims translate into pricing changes at the API layer within a couple of quarters. If they don’t, the constraint was never compute design.
- Whether the trade press and the mainstream tech press are covering the same thing. When they diverge, the trade press is usually early.
Google beating Nvidia is a good headline. It’s also the layer of this story that requires the least thought to understand, which is why it travels. The parts you can’t see are doing more work than the parts you can, and that’s true of accelerators, agent frameworks, and most of the software I get paid to be rude about.
🕒 Published: