\n\n\n\n Nobody Pitches You Their Substrate - AgntHQ \n

Nobody Pitches You Their Substrate

📖 4 min read•796 words•Updated Sep 17, 2026

Over 100 billion transistors. That’s the estimate for Meta’s third-generation MTIA accelerator, built on TSMC’s N3P with HBM memory attached. It’s a big number, and it’s the kind of number that gets top billing in every launch deck. What doesn’t get top billing is the piece of laminated resin that has to hold all of it together without warping.

I review AI tools for a living, which means I sit through a lot of hardware announcements that are functionally identical. Transistor count, memory bandwidth, some inference multiplier versus a GPU from two years ago. Cerebras says the CS-4 hits 30x faster AI inference than GPUs, with first shipments scheduled for later this quarter. That’s the headline. Nobody’s headline is the packaging.

The part of the accelerator nobody names

ABF substrate — Ajinomoto Build-up Film, named after the company that also sells you MSG — is the layered board that sits under the silicon and routes signal and power between the dies and the rest of the system. As the state of ABF substrates in data center silicon in 2026 lays out, designers are now placing multiple large logic dies alongside stacks of memory on a single accelerator package. Every one of those dies needs thousands of connections. The substrate is what makes those connections possible.

This matters because the industry has decided that the way forward is packing more silicon onto each accelerator. Not one big chip, but several large logic dies plus memory, all co-located to feed the compute with enough bandwidth to keep frontier models from starving. That approach makes the substrate bigger, more layered, and much harder to manufacture without defects. You are no longer building a small board. You are building a large one that must stay flat under thermal stress while carrying an eight-figure bill of materials on top of it.

Why a reviewer cares about a non-reviewable component

Because supply constraints on unglamorous components have a habit of deciding who ships and who slips. And 2026 has a lot of people trying to ship.

  • Cerebras raised $1B in a Series H on February 3, 2026, for its wafer-scale training and inference processor.
  • Positron AI raised $230M in a Series B on February 4, 2026, for a memory-centric inference accelerator.
  • Huawei is rolling out the Ascend 950DT.
  • Meta’s MTIA is on its third iteration.
  • Groq is pushing on inference alongside Cerebras.

Every one of those products needs advanced packaging. Every one is competing for the same specialized manufacturing capacity. When you read a launch date, you are reading a bet on that capacity being there.

The market is actually moving

NVIDIA still holds an estimated 80-85% revenue share of the data center AI accelerator market in 2026, down from roughly 92% in 2024. AMD is at an estimated 5-7%, up from about 2%. Google’s TPU is in the mix. Those are estimates, not audited figures, and I’d treat them as directional rather than precise.

Still, the direction is clear enough. NVIDIA giving up several points of share in two years is not a collapse, but it is the first real crack in a near-monopoly. Some of that is AMD executing. Some of it is hyperscalers building their own silicon and buying less of someone else’s. Some of it is a genuine market for inference-specific hardware that didn’t exist at scale before.

What this means if you’re buying compute, not chips

Most readers of this site aren’t purchasing accelerators. You’re renting inference by the token, and you care about price, latency, and whether your provider stays up. Here’s how the hardware story translates.

More vendors shipping inference silicon should mean cheaper tokens. The 30x claim from Cerebras, whatever it survives contact with reality as, is aimed squarely at inference economics. Same for Positron’s memory-centric approach and Groq’s whole thesis. Inference is where the recurring cost lives, and there is now real competition for it.

But claims like 30x are vendor claims. They are measured against a baseline the vendor picked, on a workload the vendor picked. I don’t take them at face value, and neither should you. The number that matters is what your actual model costs to serve at your actual throughput, and nobody publishes that for you.

My honest read

The interesting story in AI silicon right now isn’t the transistor count. It’s the physical limits of assembling these things. Bigger packages, more dies per package, more memory stacks, and a substrate layer that has to hold it all together at yields good enough to matter. That’s a manufacturing problem, and manufacturing problems are slow to solve and expensive to get wrong.

So when the next accelerator announcement lands with a big multiplier and a shipping quarter, my first question isn’t about the architecture. It’s whether they can actually build the package. Ask about the boring layer. That’s where the schedule slips.

🕒 Published:

📊
Written by Jake Chen

AI technology analyst covering agent platforms since 2021. Tested 40+ agent frameworks. Regular contributor to AI industry publications.

Learn more →
Browse Topics: Advanced AI Agents | Advanced Techniques | AI Agent Basics | AI Agent Tools | AI Agent Tutorials
Scroll to Top