
Share
OpenAI's new Sol and Luna models cut prices roughly in half versus GPT-5.6, reshaping the cost efficiency frontier, but the gains come with a mixed bag of benchmark improvements and some notable regressions in knowledge work tasks.
OpenAI just made its Sol and Luna models a lot cheaper to run, without meaningfully moving the needle on raw intelligence. GPT-6 Sol drops from $4/$20 to $2/$10 per million input/output tokens, and Luna falls from $0.20/$1.20 to $0.10/$0.50. The usual cache discounts carry over: 90% off for cache reads, 25% premium for cache writes. That's a straightforward price cut, not a repricing tied to some architectural overhaul, and it's the headline story here.
For practitioners, cost per task is often the number that actually matters, not just per-token pricing. On that front, the improvement is real. GPT-6 Sol (max) now costs $1.06 per task to run the Artificial Analysis Intelligence Index, down from $1.99 for GPT-5.6 Sol (max), roughly a 50% reduction. Luna (max) costs $0.07 per task versus $0.18 previously, about 60% cheaper. Worth noting: both models actually use more output tokens per task than their predecessors (31k vs 29k for Sol, 51k vs 41k for Luna). The savings are coming entirely from the price cut, not from the models becoming more token-efficient. That's an important distinction if you're trying to forecast future cost trends rather than just react to this one.
These two releases apparently let OpenAI claim a significant chunk of the cost efficiency Pareto frontier, the curve mapping the best available tradeoffs between capability and price. If you're choosing a model based on dollars-per-unit-of-intelligence, Sol and Luna just got a lot more competitive.
Coding performance splits down the middle between the two model sizes. In OpenAI's Codex harness, GPT-6 Sol (max) scores 57 on the Artificial Analysis Coding Agent Index, up 2 points from GPT-5.6 Sol (max). The gains show up in Terminal-Bench 4.0 (43% vs 37%) and SWE-Atlas-QnA (58% vs 54%). At $2.99 per task, that's roughly half the cost of the previous generation, and it lands Sol on the Pareto frontier for Coding Agent Index versus cost per task. That's a genuinely good outcome: better score, half the price.
Luna didn't get the same treatment. GPT-6 Luna (max) scores 41 on the Coding Agent Index, down 2 points from its predecessor, with drops in SWE-Atlas-QnA (44% vs 49%) and DeepSWE v1.1 (64% vs 66%). It's still roughly 60% cheaper per task, so the value proposition holds if you're cost-constrained, but if you needed Luna specifically for coding agent work, expect a small step backward in raw capability.
Hallucination rates tell a more consistently positive story. On AA-Omniscience, Artificial Analysis's knowledge and hallucination benchmark, both models improve substantially. GPT-6 Sol (max) cuts its hallucination rate from 92% to 60%. Luna goes from 93% to 77%. Sol gets there partly by being more conservative: it attempts only 83% of questions versus 99% for GPT-5.6 Sol (max). Declining to answer more often cuts wrong answers by about a quarter, but it also costs Sol 5 points of accuracy, dropping from 59% to 54%. Luna's accuracy barely moves, 44% versus 43%, while it also answers fewer questions than before. Net effect on the AA-Omniscience Index: Sol climbs from 22 to 27, and Luna jumps from -10 to 1.

That tradeoff, between answering confidently and answering correctly, is worth sitting with. A model that says "I don't know" more often will look better on hallucination metrics almost by construction. Whether that's actually more useful depends entirely on your application. For anything where a wrong answer is costlier than no answer, Sol's new behavior is a win. For tasks where you need an answer regardless, the accuracy dip is a real cost.
Other evals show a similar mixed pattern. Both models improve on AutomationBench-AA (Sol 62% vs 60%, Luna 53% vs 50%) and Terminal-Bench 4.0 (Sol 44% vs 40%, Luna 13% vs 12%). But two knowledge-work benchmarks regress noticeably. In GDPval-AA v2.1, adapted from OpenAI's dataset of economically valuable tasks spanning 44 occupations, Sol drops roughly 100 Elo points and Luna drops about 75. Luna additionally loses around 45 Elo points on AA-Briefcase v1.1, a private eval built around multi-week knowledge work projects involving thousands of input files. Sol stays level there.
The Artificial Analysis team dug into what's actually causing the knowledge-work regressions by manually inspecting hundreds of outputs. The pattern they found: shorter deliverables that skip elements the grading rubric expects, plus a general drop in presentation quality. That's not a subtle benchmark quirk, it's a real behavioral shift in how the models produce longer-form work product. If your use case involves generating polished, complete documents rather than quick Q&A or code snippets, this is the regression to watch for.
GPT-6 Sol and Luna represent a pure cost play more than a capability leap. Prices are roughly halved across both models, and that's driving real reductions in cost per task despite both models actually consuming more output tokens than their predecessors. Sol improves modestly on coding benchmarks and lands on the Pareto frontier for cost-efficient coding performance, while Luna takes a small step back on the same coding metrics.
Hallucination rates drop meaningfully for both models, though Sol's improvement comes partly from answering fewer questions rather than pure accuracy gains. The clearest warning sign is the regression on GDPval-AA v2.1 and AA-Briefcase v1.1, where both models produce shorter, less complete outputs on complex knowledge-work tasks. If your workload leans on long-form deliverables rather than coding or quick factual queries, that regression is worth testing against your own use case before you assume the price cut is a pure upgrade.
Tags
Original Sources
GPT-6 Sol and Luna push the cost efficiency frontier
↗ https://artificialanalysis.ai/articles/gpt-6-sol-and-luna-push-the-cost-efficiency-frontier
About the author
Kai built ML infrastructure at a Bay Area startup before developing an obsession with transformer architectures and inference optimisation that eventually pulled him out of product work entirely. A stint at a compute research lab sharpened his instinct for what actually matters in a model release versus what is marketing. He writes from the inside — from the perspective of someone who has debugged the systems he is describing at three in the morning. He is allergic to hype and instinctively drawn to the unglamorous plumbing questions that everyone else skips over.
More from The Engineer →This Week's Edition
23 September 2026
29 articles
Related Articles

Anthropic's Claude Opus 5.5 Ships as Five Models in One, Tuning Intelligence Against Cost
Models & Research · 5 min

OpenAI's GPT-6 Sol Lands With Six Variants and a Sharper Price-to-Intelligence Curve
Models & Research · 5 min

Heidi Overton's FDA Confirmation Hearing Arrives at a Pivotal Moment for Drug Innovation
Policy & Regulation · 5 min
Related Articles

Anthropic's Claude Opus 5.5 Ships as Five Models in One, Tuning Intelligence Against Cost
Models & Research · 5 min

OpenAI's GPT-6 Sol Lands With Six Variants and a Sharper Price-to-Intelligence Curve
Models & Research · 5 min

Heidi Overton's FDA Confirmation Hearing Arrives at a Pivotal Moment for Drug Innovation
Policy & Regulation · 5 min
More Stories
© 2026 Cedar & Bloom. All rights reserved.