A startup you have probably never heard of says it just beat Anthropic and OpenAI at one of the hardest jobs in science. That same startup has published, as far as I can find, almost nothing that would let you check the claim yourself. Sit with that contradiction for a second, because it is the whole story.
Inherent, a British AI lab founded by DeepMind alumni, announced in 2026 that its AI “teammate” — an agent called Faraday — outperformed models from Anthropic and OpenAI at replicating research. TechCrunch picked it up, the post did numbers, and now everyone wants to know whether this is a real breakthrough or another founder-flavored benchmark victory lap.
I review AI tools for a living, and I have learned to hold two thoughts at once: this could genuinely matter, and I have almost nothing to evaluate it with. Let me walk you through both.
Why research replication is actually a big deal
Replication is the unglamorous backbone of science. Someone publishes a result, and someone else has to reproduce it — same method, ideally the same outcome — before anyone should trust it. Humans are famously bad at getting around to this. It is tedious, it earns no glory, and it takes serious skill to read a paper, reconstruct the method, and run it faithfully.
So an AI agent that can do replication work well is not a toy. It is aimed at a real bottleneck. If Faraday can actually take published research and reproduce it reliably, that is a meaningful capability, and it explains why Inherent branded it a “teammate” rather than a chatbot. The framing suggests something that works alongside researchers on long, multi-step tasks rather than answering one-off questions.
And credit where due on the pedigree question: DeepMind alumni have been spinning out startups across Europe at a steady clip. These are not random founders chasing a trend. The talent is plausible. The problem choice is smart.
What Here is what it does not give us, at least in anything I have seen so far:
- What benchmark or evaluation was used, and who designed it
- Which specific Anthropic and OpenAI models Faraday was measured against
- Whether the evaluation was independently run or scored in-house
- What “replicating research” meant in practice — full experimental reproduction, or something narrower
Those gaps matter enormously. “We beat the big labs” is the single most common claim in AI startup marketing, and the details almost always decide whether it holds up. A vendor comparing its purpose-built agent against general-purpose models on a task the vendor selected is not a fair fight — it is a product demo wearing a lab coat. That does not mean Inherent did that. It means we cannot yet rule it out, and neither can you.
There is also a naming irony I refuse to leave on the table. Faraday, as in the Faraday cage — a structure famous for blocking signals from getting through. An apt mascot, given how little signal has escaped this announcement.
The pattern I keep seeing
Small labs punching up at Anthropic and OpenAI is now a genre. The playbook: pick a specific vertical task, build an agent tuned for it, run an evaluation, announce a win. Sometimes the win is real — specialization genuinely beats generality on narrow tasks. Sometimes the win evaporates the moment a third party reruns the test.
The honest position is that specialization is exactly where smaller labs should be competing. You do not outspend OpenAI on compute. You outmaneuver them on focus. Research replication is a well-chosen niche: valuable, verifiable in principle, and underserved. If I were founding an AI lab with DeepMind credentials, this is roughly the bet I would make.
My verdict, for now
Interesting claim, incomplete evidence, sensible strategy. That is where I land.
What would change my mind: a published methodology, named model versions on both sides, and ideally an independent party rerunning the evaluation. What would sink it: silence, or a benchmark that turns out to have been graded by the people who built the student.
Until then, treat “outperformed Anthropic and OpenAI” the way a good scientist treats any single unreplicated result — as a hypothesis, not a fact. Which, given what Faraday is supposedly built to do, is an irony Inherent should appreciate more than anyone. If your product’s entire pitch is that claims should be reproduced before they are believed, publish the receipts. I will be the first to update my review when they do.
🕒 Published: