\n\n\n\n Nobody Ever Benchmarked A Substrate - AgntHQ \n

Nobody Ever Benchmarked A Substrate

📖 4 min read•781 words•Updated Sep 16, 2026

When was the last time you read a spec sheet and cared about the fiberglass?

Probably never. Nobody does. We argue about transistor counts and FLOPS and whose interconnect is fastest, because those are the numbers vendors put on stage in giant fonts. Meanwhile the part that decides whether any of those numbers ship on schedule is a laminate board almost nobody in AI can name.

I’ve been reviewing AI tools long enough to notice a pattern: the thing everyone benchmarks is rarely the thing that constrains the product. In the accelerator world, the logic die gets the press tour. The memory stack and the substrate under it get the delays.

What the market share numbers actually tell you

NVIDIA sits around 80-85% of data center AI accelerator revenue in 2026, down from roughly 92% in 2024. AMD is up to about 5-7% from around 2%. Those two lines moved in the same direction at the same time, and the easy story is that AMD finally built a good chip. MI300X+ is competitive, sure. But a 7-point swing in a market this supply-constrained isn’t purely a story about who has the better shader compiler. It’s a story about who could get parts.

Look at what a modern accelerator physically is. Designers now place multiple large logic dies next to a stack of memory on a single package. That whole assembly sits on an ABF substrate, and every added die, every added memory stack, makes that substrate bigger, denser, and harder to yield. You can tape out the most beautiful compute die in the industry and still ship late because the board underneath it warped.

Cerebras and the wafer-scale dodge

This is the context that makes Cerebras interesting rather than just weird. The WSE-3 approach — one enormous wafer instead of many chiplets stitched together — is usually pitched as a bandwidth story. It’s also a packaging story. Fewer die-to-die boundaries, fewer things to assemble, a different set of manufacturing headaches instead of the industry-standard ones.

Cerebras raised $1B in a Series H on February 3, 2026. That is not a number you raise on benchmark charts alone. That is a number you raise when investors believe the standard recipe has a structural ceiling.

One day later, on February 4, Positron AI raised $230M in a Series B for a memory-centric inference accelerator. Read those two rounds back to back. Within 48 hours, the market wrote nearly $1.25B in checks to two companies whose pitch is essentially “the memory and assembly path is the problem, not the math.”

The losers make the same point

Intel’s Gaudi 3 underperforms. Broadcom’s custom XPUs are gaining traction. Both facts point the same direction, and it isn’t flattering to the conventional wisdom.

Gaudi 3 exists. It’s real silicon from a company that knows how to manufacture. It still isn’t winning, which suggests raw compute capability was never the scarce resource. Broadcom, meanwhile, is not selling the fastest general-purpose AI chip. It sells the ability to design and deliver custom accelerators for customers who already know their workload — including the packaging, memory integration, and supply chain work that makes the thing manufacturable.

Meta’s MTIA is on its third iteration, expected on TSMC’s N3P, very likely north of 100 billion transistors, with HBM memory. Note that the memory choice is part of the headline description of the chip. It isn’t a footnote. For a chip designed to run one company’s specific workloads, memory is the design.

What this means if you’re buying, not building

You aren’t going to buy a substrate. But you are going to make decisions downstream of this, and most of the standard advice is bad:

  • Availability beats specs. A slightly slower accelerator you can actually get is worth more than a faster one on a waitlist. AMD’s share gain is partly this.
  • Treat memory capacity and bandwidth as the primary number. If your workload is inference, this is not a tiebreaker. It’s the whole evaluation.
  • Be skeptical of any vendor comparison built only on peak compute. Peak compute is the easiest number to win and the least predictive of what you’ll observe.
  • Watch funding rounds as a signal, not a scoreboard. Two memory-and-packaging bets funded in two days is the market telling you where the pain is.

NVIDIA is still winning by a wide margin, and nothing here suggests that flips soon. But the erosion from 92% to the low 80s happened while the industry hit a physical wall in how much silicon fits on one package. Those aren’t unrelated events.

The chip everyone talks about is the one doing the multiplication. The chip that decides whether your cluster arrives in Q2 or Q4 is the one holding everything together, and it doesn’t have a keynote.

🕒 Published:

📊
Written by Jake Chen

AI technology analyst covering agent platforms since 2021. Tested 40+ agent frameworks. Regular contributor to AI industry publications.

Learn more →
Browse Topics: Advanced AI Agents | Advanced Techniques | AI Agent Basics | AI Agent Tools | AI Agent Tutorials
Scroll to Top