Remember when data labeling was the least glamorous job in AI? Back when chatbots were soaking up every headline, the companies quietly paying humans to draw boxes around stop signs and rank model outputs were treated like plumbing. Necessary, unsexy, and definitely not where the money was. Then the models got good, and everyone figured out that the plumbing was actually the product.
We’re watching the sequel now. Mecka AI is nearing a $500 million valuation in a round led by Sequoia Capital, and the driver isn’t a model, an agent, or a chip. It’s data. Specifically, the kind of data robots need to learn how to move through the physical world, grasp objects, and finish tasks that a toddler handles without thinking.
Why this round is different from the last hundred
I review AI tools for a living, which means I spend most of my week watching companies raise enormous sums on the promise that a wrapper around someone else’s model constitutes a business. This one is a different shape. The bet here isn’t that Mecka has a smarter architecture. It’s that physical-world training data is scarce, expensive to produce, and cannot be scraped.
That last part matters more than anything else in this story. Language models got their start by hoovering up the internet, essentially for free. There is no internet of robot motion. Nobody has a trillion-token corpus of hands picking up unfamiliar objects at odd angles in badly lit kitchens. If you want that data, someone has to physically generate it, over and over, and pay for the labor and hardware to do it. Scarcity plus necessity is how you get a valuation like this attached to what is, functionally, a data operation.
The numbers we actually have
Reporting notes that as of early June, Mecka was projecting it would finish 2026 at an annual run rate of $100 million, according to comments its Gao made to Fortune around a previous fundraise. That’s a projection, not a result, and I’d treat it accordingly. But even taken at face value, a $500 million valuation against a $100 million projected run rate is a comparatively grounded multiple by 2026 standards. I’ve seen agent startups with a demo video and a waitlist ask for more.
What I don’t have, and what nobody outside the round has, is the interesting stuff:
- How much of that revenue comes from a small number of large robotics customers
- Whether the data is exclusive or gets resold to competing labs
- How the collection cost scales as task complexity increases
- What happens to demand if simulation-generated data gets good enough to substitute
That fourth one is the actual risk. Real-world data collection is expensive and slow. Synthetic and simulated data is cheap and fast, and it keeps getting less terrible. Every dollar of Mecka’s valuation is a bet that the reality gap stays wide enough that real-world demonstrations remain necessary. That’s a defensible bet today. It’s not a permanent one.
What this signals about where AI money is going
The center of gravity in AI investment is moving a step past large language models. Not away from them, past them. Text is close to saturated as a frontier. Physical manipulation is wide open, and the constraint isn’t compute or model design, it’s examples. Whoever supplies the examples sits in a good position regardless of which robotics company wins.
That’s the strategic logic Sequoia is buying. Picks and shovels, again. It worked for the labeling companies in the LLM era and it’s a reasonable pattern to run twice. My skepticism isn’t about the thesis, it’s about the pricing. Half a billion dollars on projected revenue in a category where the technical substitute is improving fast is a bet with real timing risk baked in.
My honest read
I can’t review Mecka AI as a tool, because it isn’t one. You don’t sign up for it. There’s no interface for me to tear apart, no pricing page to complain about, no onboarding flow that wastes ten minutes of your life. This is infrastructure, and the customers are companies building robots, not people reading tool reviews.
So take this as a market signal rather than a recommendation. If you’re building anything involving physical automation, the cost and availability of training data is about to become one of your defining constraints, and it’s being priced accordingly right now. If you’re an investor, you’re being asked to believe that real-world robot data stays hard to get for several more years.
The part I’ll be watching is whether the run-rate projection holds and whether the customer base broadens beyond a handful of well-funded robotics labs. Data businesses look excellent when demand is a land rush and considerably worse when the rush ends. Right now, it’s a rush.
🕒 Published: