Quick question. When you picture an AI accelerator, what do you see? Probably a slab of silicon with a fan on it, some benchmark chart, a CEO holding it up like a newborn. Now tell me what it’s sitting on. Not the server. Not the rack. The thing directly underneath the die.
Most people reviewing AI hardware can’t answer that, and I include a fair number of analysts who get paid to. The answer is a substrate, usually built on ABF material, and it is quietly one of the most important components in the entire stack. It has no brand loyalty, no launch event, no subreddit. It just decides whether your accelerator ships.
Why the boring layer suddenly matters
Here’s what changed. The industry stopped being able to win purely by shrinking transistors, so it started winning by cramming more silicon into one package. Multiple large logic dies sitting side by side. Stacks of high-bandwidth memory crowded around them. Microsoft’s Maia 200, unveiled in January 2026 on TSMC’s 3nm process, carries 216GB of HBM3e. That is not a chip in the way people used to mean the word. That is a small city of chips wearing one heatspreader.
All of that has to be mounted on something, connected to something, and kept flat while it heats up and cools down thousands of times. Bigger packages mean bigger substrates. Bigger substrates mean more layers, tighter tolerances, and lower yields. The material stops being a commodity and starts being a constraint.
The reporting on ABF substrates in 2026 makes the point plainly: designers are packing ever more silicon onto each accelerator to hit the compute and memory bandwidth frontier models demand. The packaging is where that ambition meets physics.
What this means when you read a spec sheet
I review AI tools and the infrastructure under them, and my running complaint is that spec sheets are written to flatter, not inform. Peak FLOPS is the number everyone quotes. It is also the number least likely to survive contact with a real workload.
Memory bandwidth and how memory is physically attached tell you far more about how a chip behaves when you actually use it. That is why the funding pattern is interesting. Positron AI raised $230M in a Series B in February 2026 for a memory-centric inference accelerator. The name of the category says the quiet part: the bottleneck moved. Cerebras closed a $1B Series H a day earlier on a wafer-scale approach, which is a different bet on the same problem. If you can’t get data to the compute fast enough, make the compute and the memory neighbors instead of pen pals.
Neither of those approaches works without packaging that can physically hold it together. Wafer-scale, in particular, is a packaging thesis dressed as a compute thesis.
The supply chain nobody wants to talk about
Nvidia, AMD, and Broadcom lead this market in 2026, and all three are pouring money into new technology and infrastructure. Beneath them sits a layer of startups going after inference and custom silicon. Above them sit the cloud providers building their own parts, with TrendForce projecting custom ASIC shipments from those providers to grow 44.6% in 2026.
Notice what everyone in that list has in common. They all need advanced packaging, and they’re all drawing from the same narrow set of suppliers to get it. When demand for compute grows faster than the ability to build the thing the compute sits on, the constraint stops being design talent and starts being industrial capacity.
My read, and this is opinion rather than reported fact: the most interesting competitive question in AI hardware right now isn’t whose architecture is cleverest. It’s who has locked in packaging and memory supply. Architecture can be copied in a design cycle. A supply agreement can’t.
How to be a less gullible reader
A few habits I’d suggest when the next accelerator announcement lands:
- Look for memory capacity and bandwidth before you look at compute throughput. It’s the number that predicts real inference behavior.
- Ask whether the part is in mass production or merely announced. Maia 200 is described as in mass production and serving Microsoft 365 Copilot, which is a materially different claim from a slide.
- Treat package size and die count as a signal about cost and supply risk, not just performance.
- When a vendor talks only about peak performance, assume the other numbers were less flattering.
None of this makes for good keynote theater. There’s no dramatic reveal in a laminate. But the difference between a chip that exists on a roadmap and one that exists in a data center often comes down to whether the unglamorous layer underneath it could be built at volume.
Next time someone holds up a piece of silicon on stage, look at the part they’re holding it by. That’s where the real story lives, and nobody’s going to tell it to you.
🕒 Published: