It’s 11:40 on a Tuesday morning. Your support agent has been running fine for six weeks. Then a customer pastes in an order number and the agent replies with a summary of a completely different account. You check the logs. Nothing. No error, no timeout, no rate limit, no failed tool call. The trace looks identical to the four hundred that worked. You retry the same input and it behaves perfectly. You close the ticket with a note that says “could not reproduce,” and you go get coffee.
That coffee is the problem. Not the bug. The coffee.
We stopped demanding explanations
A post on i hate the future in late September 2026 put a name to something a lot of us have been quietly living with: the normalization of inexplicable failures. Its framing stuck with me. A broken door is comic. You push, it doesn’t open, everyone laughs, you find another door. A broken AI service is something else entirely, because it makes a business process impossible to trust. The writer’s line about software engineering today is that we’re actively building systems whose failures we cannot account for.
The piece hit the front page of Hacker News, which tells you the nerve was already exposed. Engineers didn’t need convincing. They’ve been shipping “could not reproduce” as a resolution status for two years now.
I review AI tools for a living. I run agents against the same inputs dozens of times. And the single most common pattern I see in vendor behavior isn’t a lie or a hallucinated benchmark. It’s a shrug. You report a failure, you get told to adjust your prompt. You report it again with a reproduction, you get told the model was updated. You ask which version you were running and when it changed, and the conversation ends.
The vocabulary of not knowing
An entire dialect has grown up around this. Watch for it in docs, changelogs, and support replies:
- “Non-deterministic by design” — a description of a property, repurposed as an excuse for an unbounded failure mode.
- “Try rephrasing” — the debugging advice equivalent of percussive maintenance.
- “Improved model behavior” in a changelog, with no note on what behavior changed or what broke.
- “Edge case” — applied to inputs that make up a meaningful share of real traffic.
- “Working as intended” — when nobody can state the intent precisely enough to test it.
None of that would survive a code review in a payments system. In an agent product it passes as documentation.
The doom argument is a distraction from the boring one
Jacob Coxen, a young researcher who worked at Anthropic, has been warning that AI and superintelligence could kill everyone within ten years. That’s a real warning from someone with real proximity to the work, and 2026 has been full of that genre: existential warnings on one side, critiques of software engineering practice on the other.
My read is that the two are the same story told at different volumes. If we can’t explain why an agent misrouted an order number on a Tuesday, the claim that we’ll explain the behavior of something far more capable is a fantasy. The extinction debate gets the headlines. The unglamorous version — nobody can say why the thing did what it did, and we shipped it anyway — is the load-bearing premise underneath it. You don’t have to buy the ten-year timeline to be bothered by the fact that our floor for “explainable” has dropped through the basement.
What I now hold vendors to
I’ve changed how I score tools because of this. Accuracy numbers matter less to me than failure legibility. Concretely, what I want from any agent product:
- Pinned model versions, with dates, and a changelog that says what moved.
- Traces that show the actual inputs and outputs of every tool call, not a summary of intent.
- A stated confidence or refusal path, so the system can say “I don’t know” instead of inventing an account.
- Reproducibility controls — a seed, a temperature, something — so a bug report can be a bug report.
- A support process that treats “could not reproduce” as an open investigation, not a closed ticket.
Tools that can’t offer these aren’t unreliable because AI is hard. They’re unreliable because nobody built the instrumentation, and the market let them get away with it.
Stop normalizing it
The broken door analogy works because a door has exactly one job and you can see it fail. Your agent has hundreds of jobs and fails invisibly, in ways that look like success in your dashboard. That’s not a more advanced kind of software. It’s a less accountable kind.
So the next time your agent does something you can’t explain, don’t close the ticket. File it with the vendor, ask for the version, ask for the trace, and ask them to explain it. The acceptance is the failure mode. Everything else is just a bug.
🕒 Published: