Jake Handy’s writeup of the launch skipped the marketing and went straight to the spec sheet: 1.05M-token context window, 922K of that available for input, 128K max output, identical across both models, with knowledge cutoffs that differ by a month. That is the correct instinct. When a lab ships two models on the same day with the same context budget and the same shape, the interesting story is never the model. It’s the invoice.
On September 22, 2026, OpenAI dropped GPT-6 Sol and GPT-6 Luna as direct replacements for the GPT-5.6 lineup, with prices cut 50% across the board. Sol lands at $2 per million input tokens and $10 per million output. Luna comes in at $0.10 per million input and $0.50 per million output. Both are positioned as the cheaper, faster siblings to GPT-6 Astra, the flagship that shipped earlier the same month.
What the pricing actually tells you
A 50% cut is not a courtesy. Labs don’t halve prices on models people are happily paying full rate for. They halve prices when the cost of serving has dropped far enough that the old number looks indefensible, or when somebody else’s number is making their number look bad. The coverage around this launch has been framing it as undercutting Claude, and that framing is doing most of the work in the headlines.
Look at the gap between the two models. Sol input costs 20x what Luna input costs. Output is 20x as well. That is a wide spread for two models sharing a context window and shipping on the same day. It tells you OpenAI has stopped pretending there’s one right model for every job and started pricing tiers the way cloud providers price instance types. You pick based on how much the answer is worth, not based on which model is “best.”
For anyone running agents in production, this is the number that matters. Agent loops are output-heavy and repetitive. A retry-happy agent chewing through tool calls at $10 per million output tokens is a different business than the same agent at $0.50. That 20x is the difference between a feature you can ship to free users and a feature you gate behind a paid tier.
The cutoff dates are the weird part
Sol’s knowledge cutoff is April 20, 2026. Luna’s is May 18, 2026. The cheaper model knows about a month more of the world than the expensive one.
I don’t want to over-read this, but it’s worth sitting with. The conventional assumption is that the pricier model is the more recent, more capable artifact and the cheap one is a distilled leftover. Here the cheap one has fresher data. Either the two models came out of different training runs on different schedules, or Luna was finished later and Sol was already locked. Either way, if your use case touches anything that happened in late April or early May of 2026, the $0.10 model may serve you better than the $2 one. That is not a sentence I expected to write.
The context window math
922K input plus 128K output equals 1.05M. The window is shared, not additive. If you stuff 900K tokens of context in, you are not getting 128K of output on top of a fresh budget. That’s standard, but I’ve watched enough teams build retrieval pipelines that assume otherwise to flag it.
The more practical point: a million-token window on a $0.10-per-million-input model changes what’s economically reasonable to send. Filling Luna’s input window completely costs roughly nine cents. Filling Sol’s costs about $1.84. When stuffing context is nearly free, the calculus on whether to build careful retrieval versus just throwing the whole document at the model shifts. Not because retrieval stopped being the better engineering choice, but because the lazy option stopped being expensive enough to punish you.
What I’d actually do
- If you’re on GPT-5.6 Sol, the migration is a price cut with a name change. Test it, but the direction is obvious.
- Route classification, extraction, summarization, and high-volume agent steps to Luna. At $0.10 input, you can afford to be wasteful.
- Reserve Sol for steps where a wrong answer costs real money. The 20x premium needs to buy you something measurable, and you should be the one measuring it.
- Check your cutoff dependency before assuming Sol is the safer pick. It isn’t always.
OpenAI has not made a case here about capability, and I’m not going to invent one. What they’ve made is a case about cost, aimed squarely at teams doing arithmetic on their inference bills. On that narrow question, the launch is persuasive. Everything else is still unproven until someone runs their own evals, which is exactly the work nobody wants to do and everybody should.
🕒 Published: