Alibaba gives away its best models for people to download and run for free. Alibaba also wants $100 billion in combined cloud and AI revenue by 2031. Hold those two facts next to each other and you have the most interesting strategic puzzle in AI right now.
Most companies pick a lane. You either sell the model as a product, the way OpenAI and Anthropic do, or you sell the hardware, the way Nvidia does, or you rent the compute, the way the hyperscalers do. Alibaba has decided the answer is all of it. Own the chips. Own the CPUs. Own the cloud that runs them. Own the models on top. Then release the model weights so widely that the rest of the industry builds on your architecture whether or not they ever pay you a cent.
What was actually announced
The May 20, 2026 push was the loudest part: a new flagship large language model, a homegrown AI chip the company says triples the performance of its predecessor, and a rebuilt cloud underneath both. In September, Alibaba filled in more detail on the Qwen roadmap and the full-stack infrastructure, and the stock went up on it. The company also landed on TIME’s 2026 list of most influential companies, largely on the strength of its open-model position.
The current flagship story runs through Qwen 3.7-Max, built specifically for agentic workloads, with Qwen 3.8 Max arriving in August 2026 and Alibaba claiming it can rival Anthropic’s best work. If you find that release cadence dizzying, you are reading it correctly. Point-one version bumps every few months is not a roadmap, it is a treadmill.
Why agentic workloads change the hardware math
This is the part I think gets underplayed. When Alibaba says a model is designed for agentic workloads, that is not marketing garnish. Agents behave nothing like chatbots at the infrastructure level. A chat turn is one request, a few thousand tokens, done. An agent loops. It calls tools, reads results, re-plans, calls again, and burns through context windows in a way that turns a single user task into dozens or hundreds of inference passes.
That shifts the bottleneck. Serving agents economically is less about peak benchmark scores and more about cost per token at sustained volume, memory bandwidth, and how cheaply you can keep long contexts warm. Which is exactly the kind of problem you solve by controlling your own silicon rather than queuing behind everyone else for someone else’s accelerators.
So the chip investment and the agentic model design are the same bet wearing two hats. If agents become the dominant shape of AI workloads, whoever owns the cheapest inference stack wins the volume, and the model quality only has to be good enough to stay in the conversation.
The open-model contradiction, resolved
Giving models away looks like leaving money on the table until you think about where the money actually is. Open weights do a few things for Alibaba that a paid API never could:
- They make Qwen the default starting point for developers and startups who were never going to pay for a frontier API anyway.
- They build a fine-tuning and tooling ecosystem around Alibaba’s architecture rather than someone else’s.
- They create a natural funnel, because the easiest place to run a Qwen model in production at scale is the cloud built specifically to run it.
The models are the distribution. The infrastructure is the business. It is a well-worn playbook, just executed at a scale most companies could not attempt.
Where I stay skeptical
Full-stack is a lovely word for investors and a brutal one for engineers. Every layer you own is a layer you have to keep competitive on your own, forever. Chips are the hard part. A tripling over your own previous generation is a real achievement and also a comparison against yourself, which is the easiest benchmark in existence.
The $100 billion target by 2031 is the number I would keep an eye on, because it is the one thing here that cannot be spun. Model launches are announcements. Chip performance claims are self-reported. Revenue is revenue. Five years is long enough for the agentic thesis to either pay off spectacularly or look like a very costly detour.
What this means if you build with agents
Practically speaking, you now have a serious open-weight option built specifically for the workload you care about, backed by a company with the capital to keep iterating. That is good for you regardless of how the corporate strategy plays out. Test Qwen against whatever you are paying for now on your own agent traces, not on benchmarks. Watch inference cost per completed task, not per token.
The vertical integration story is Alibaba’s problem. The cheaper inference is your opportunity.
🕒 Published: