
Share
OpenAI's latest release splits into five reasoning-effort variants with a 5.5x price spread. Here's what the benchmarks actually say about the tradeoffs between intelligence, speed, and cost.
OpenAI quietly dropped GPT-6.1 Sol in September 2026, and instead of one model, you get five: low, medium, high, xhigh, and max reasoning variants, all built on the same underlying architecture but tuned for different points on the intelligence-versus-cost curve. If you've been tracking OpenAI's release cadence, this is now the standard playbook. Ship a family, let developers pick their tradeoff.
The headline number is 52 on the Artificial Analysis Intelligence Index, achieved by the max variant. That's a modest step up from GPT-6 Sol's 48 but still trails GPT-6 Astra's 53, which remains OpenAI's highest-scoring model to date. So this isn't a frontier-pushing release so much as an efficiency and product refresh, more reasoning tiers, better price-to-performance at the low end, same architecture family.
Here's the breakdown across the five variants:
Prices swing up to 5.5x between the cheapest and most expensive tier for a single Intelligence Index task. That's a meaningful spread if you're running high-volume production workloads where cost per call compounds fast.
The Intelligence Index score isn't a single test, it's a weighted composite. Artificial Analysis built this version (v4.3.2) from ten separate evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1. That's a mix of coding benchmarks (Terminal-Bench, SciCode), knowledge-heavy tests (Humanity's Last Exam, AA-Omniscience), and business-task simulations (GDPval, AA-Briefcase) designed to approximate real economic value rather than just academic puzzle-solving.
Worth noting: this methodology has shifted versions before (previous Sol and Astra releases used earlier index versions), so cross-generation comparisons carry some noise. Still, within this snapshot, GPT-6.1 Sol max sits in a competitive middle tier among the 686 models Artificial Analysis currently tracks, ahead of GPT-6 Sol and GPT-5.6 Terra, behind GPT-6 Astra.

There's also a set of domain-specific capability indices worth checking if you're evaluating for a vertical use case, Finance & Accounting, Legal, Healthcare & Medical, Engineering, Economics, and Strategy & Ops. Each pulls a different subset of the ten core evaluations. The Engineering Index, for instance, leans on CritPt and Terminal-Bench alongside Omniscience and HLE, which makes sense if you're specifically evaluating for coding-agent or technical-reasoning workloads rather than general knowledge tasks.
On the architecture side, all five variants ship with a 1M token context window, roughly 1,500 pages of standard document text, and support text-and-image input with text-only output. No parameter count is disclosed, which has become the norm for frontier releases from OpenAI. Weights aren't available either; this is a proprietary, API-only release, distributed exclusively through OpenAI's own platform for now.
One thing that stands out: GPT-6.1 Sol (max) hits both the highest intelligence score and the highest output speed simultaneously. Normally you'd expect an inverse relationship, more reasoning tokens means slower generation. That the top-intelligence variant is also the fastest suggests some genuine efficiency work went into the reasoning pipeline itself, not just a dial turned up on compute.
Pricing sits at $1.5 per million tokens across the board according to the comparison table, which is the raw token price, distinct from the per-task cost figures above that account for the actual token volume each benchmark task consumes (input, reasoning, cache hits and writes, and answer tokens all get weighted differently). That distinction matters if you're doing your own cost modeling: a flat per-token price doesn't tell you much until you know how many reasoning tokens a given task actually burns through.
If you're picking a variant for production, the choice mostly comes down to what you're optimizing for. Low latency and lowest cost point you toward the low or medium tiers, accepting a meaningfully lower intelligence score (42-48 versus 52) in exchange. If raw capability matters more than cost, max is the obvious pick, and the fact that it's also the fastest variant removes what would otherwise be an awkward tradeoff.
Compared to the rest of the GPT-6 family, this release reads as an iteration rather than a leap. GPT-6 Astra still holds the intelligence crown at 53, and GPT-6.1 Sol's real contribution seems to be tightening the price-performance curve at the lower end of the reasoning-effort spectrum, plus that unusual speed-at-max-intelligence result. Worth watching whether OpenAI folds these efficiency gains into whatever comes after Astra, or whether Sol stays a separate, cost-optimized track.
Tags
Original Sources
GPT-6.1 Sol Models - Intelligence, Performance & Price Comparison | Artificial Analysis
↗ https://artificialanalysis.ai/models/releases/gpt-6-1-sol
About the author
Kai built ML infrastructure at a Bay Area startup before developing an obsession with transformer architectures and inference optimisation that eventually pulled him out of product work entirely. A stint at a compute research lab sharpened his instinct for what actually matters in a model release versus what is marketing. He writes from the inside — from the perspective of someone who has debugged the systems he is describing at three in the morning. He is allergic to hype and instinctively drawn to the unglamorous plumbing questions that everyone else skips over.
More from The Engineer →This Week's Edition
30 September 2026
28 articles
Related Articles

What the Uncanny Valley Actually Tells Us About Machine Perception
Models & Research · 5 min

The Chip That Let Hardware Be Rewritten Like Software: Xilinx's XC2064
Models & Research · 5 min

New Graphite Study Finds AI Models Still Have Distinctive Writing Tells, Even After Scrubbing Em-Dashes
Models & Research · 5 min
Related Articles

What the Uncanny Valley Actually Tells Us About Machine Perception
Models & Research · 5 min

The Chip That Let Hardware Be Rewritten Like Software: Xilinx's XC2064
Models & Research · 5 min

New Graphite Study Finds AI Models Still Have Distinctive Writing Tells, Even After Scrubbing Em-Dashes
Models & Research · 5 min
More Stories
© 2026 Cedar & Bloom. All rights reserved.