Picture a German radio operator in 1918, hunched over a set, tapping out a warning about enemy warship movements. The message goes out encrypted. Whoever was supposed to read it either read it or didn’t, and then the war ended, the operator went home or didn’t, and the message sat there. For 108 years it stayed unreadable. Not because nobody tried — cipher hobbyists and historians have chewed on lists of unsolved wartime cryptograms for decades — but because the thing simply would not open.
In 2026 it opened. GPT-6 Astra cracked it. Inside was intelligence about British warship movements, and the decryption was checked against surviving historical naval logs.
I review AI tools for a living, which mostly means watching demos that fall apart the second you change one variable. So let me be clear about why this particular result caught my attention, and why it’s still smaller than the headlines want it to be.
This is the rare AI result that can’t be fudged
Almost every claim I get pitched is unfalsifiable in practice. A model scores higher on a benchmark that the vendor also influences. An agent completes a task in a recorded demo, and you have no idea how many takes it took. A coding assistant “improves developer productivity by 40%,” measured by a survey the vendor wrote.
A cipher is not like that. A cipher has exactly one right answer and it existed before the model did. Either the plaintext comes out as coherent German that matches known events, or it comes out as garbage that someone has to spin. There’s no partial credit, no rubric, no judge model grading the output with a generous curve.
And this one had an external check: historical naval logs. The decrypted message described warship movements, and those movements are on record in documents nobody involved in the AI industry produced. That’s the part that matters. The model didn’t just generate something plausible. It generated something that agrees with an independent paper trail from 1918.
If you want a mental model for evaluating AI claims, use that one. Ask what the independent paper trail is. Most of the time there isn’t one.
What this does not prove
Now the cold water, because the coverage around this event got loose fast.
A WWI-era hand cipher is not modern cryptography. It’s not close. These systems were built to be operated by a human under time pressure with a codebook and a pencil, and their security assumptions died roughly a century ago. Breaking one is a hard, genuinely impressive pattern-recognition and language problem. It is not evidence that anything protecting your bank account, your messages, or your crypto wallet is in danger. I saw at least one outlet frame an AI cipher result with the question “Is Bitcoin next?” That’s not analysis, that’s traffic bait. The math involved is a different universe.
Second: one solved cipher is one data point. It tells you the system can do this, once, on this artifact, under whatever conditions the people running it set up. It doesn’t tell you the success rate across the rest of the unsolved pile. Unsolved ciphers are a graveyard of near-misses, and the ones still standing tend to be standing because of missing keys, damaged transmissions, or garbled transcription rather than pure difficulty. If the next twenty attempts fail, that’s normal, and it won’t get the same coverage.
Third, and this is the part that always gets edited out of the press cycle: humans did a lot of the work. Somebody located and transcribed the message. Somebody knew enough about 1918 German naval communications to constrain the problem. Somebody pulled the naval logs and confirmed the plaintext lined up with real fleet movements. The model was the engine. It wasn’t the whole vehicle, and the historians in this chain deserve the byline as much as the compute does.
Why I’d still put this in the win column
Because it’s a real task with a real answer that resisted real experts for over a century, and the result held up to outside verification. That combination is rare enough in AI that I’ll take it seriously even while I pick at the framing around it.
It also points at where these systems are actually useful right now, which is not “replace the specialist.” It’s “give the specialist a tool that can brute-search a space of hypotheses faster than a human can, then let the human check it against reality.” Archives, historical linguistics, damaged manuscripts, undocumented file formats — there’s a large pile of legitimately stuck problems that look like this, and they’ve been stuck mostly for lack of patient search capacity.
My advice for reading the next wave of these stories: skip the adjectives and go straight to the verification. Ask what independent record confirmed the output. If the answer is “another model checked it” or “the researchers found it convincing,” you’re looking at a press release. If the answer is a stack of century-old logs written by people who had no idea any of this was coming, pay attention.
This one had the logs.
Editorial note for the desk: the brief lists Amazon as the developer of GPT-6 Astra, but every source snippet supplied with it attributes GPT-6 Astra to OpenAI. I left the developer unnamed in the copy rather than publish a contested attribution. Worth resolving before this goes live.
🕒 Published: