
Share
A four-month capability gap and a five-fold price differential are pushing companies toward Chinese open models for routine work, leaving premium closed systems to justify their cost only on the hardest tasks.
The math on AI spending is shifting fast. According to a new report from Mozilla, the performance gap between US frontier AI models and the best Chinese open-weights alternatives has narrowed to just 4.4 months. That is a thin margin for products that often cost five times more per task.
The finding comes from Mozilla's latest State of Open Source AI report, published September 15 and shared with Ars Technica ahead of release. Its central claim: most organizations should now default to open models for the bulk of their AI workloads, reserving closed frontier systems for a narrow set of high-value tasks.
The numbers back up the thesis. Moonshot AI's Kimi K3, an open model, scores just three points behind Anthropic's closed Fable 5 on the Artificial Analysis Intelligence Index. It does so at roughly 30 percent of the cost. Separately, when benchmarking firm Vals AI ran Z.ai's GLM 5.2 against Anthropic's Claude Opus 4.7 and 4.8 on a neutral harness in the Terminal-Bench 2.1 evaluation, GLM 5.2 landed within a single point, at about one-fifth the per-task price.
"A closed model earns its premium in a few places: expert professional work, high-intensity retrieval, and long context," said Raffi Krikorian, Mozilla's chief technology officer, in an email to Ars Technica. His framing matters for procurement teams: the decision to pay up is workload-specific, not organization-wide.
Mozilla measured the gap using a time-horizon framework developed by the research nonprofit METR, which defines a model's capability by the length of task, measured in human-expert hours, that it can complete reliably 50 percent of the time. That horizon has been doubling on an accelerating cadence.
Right now, the best closed model can handle a task 1.7 times longer than the best open model can reliably finish. Krikorian put it plainly: "If the open frontier can handle a seven-hour job, the closed frontier can handle a 12-hour one. In four months, the open model handles the 12-hour job, and the closed one handles something around 20."
That leaves a specific window, roughly eight to 12 hours of task length, where closed models still hold a clear edge. Below that, either model type will do, and the cheaper option wins by default. Above it, neither model is reliable yet.
Krikorian's advice to buyers is blunt: "Pay when that head start is worth it, something like a deadline that lands before the open frontier catches up would be here. Routine work you'll still be doing next quarter is not, because you'll be able to do it for a fifth of the cost soon, and the model won't be the bottleneck anyway."

Real-world adoption already reflects this logic. Delivery company DoorDash has reportedly been running Kimi for routine work while routing harder tasks, the kind that would otherwise eat into expert staff time, to Fable. On OpenRouter, the AI gateway that aggregates access to hundreds of models, eight of the top 10 models by token volume in August 2026 were open-weights systems.
Caveats apply. Closed frontier models frequently ship with proprietary harnesses, the software layer that lets a model use tools and memory to act autonomously. A model tuned to its own harness can outperform the same model bolted onto third-party infrastructure, which complicates apples-to-apples comparisons. Vals AI's neutral-harness approach is one attempt to control for that variable, and it is notable that the gap held up even under those conditions.
Revenue tells a different story than usage, at least for now. A paper by Frank Nagle and Daniel Yue for the Linux Foundation found open models captured just 4 percent of overall AI revenue between May and September 2025, with closed models taking the remaining 96 percent. Krikorian expects that split has shifted materially since then, given the surge in open-model deployment over the past year, but the lag between usage and monetization is itself a signal worth watching.
There is a geopolitical dimension too. "Most of the open models the world runs on are Chinese," Krikorian said. "The Chinese labs are running the same playbook the Americans ran with Android, give it away, but own the ecosystem around it."
His concern is not nationality so much as concentration. The best closed models sit with US companies. The best open models sit with Chinese labs. Neither pattern is healthy on its own. "The uncomfortable truth is that the plural ecosystem around open-weight AI is largely funded by Chinese capital right now," he said. "That's a plurality and a concentration at the same time."
Krikorian points to the history of Linux and open internet protocols as a possible template for a Western counterweight, one built by "institutions with a mission rather than a market" rather than frontier labs chasing revenue. He specifically cites Switzerland's national compute initiative, which produced the Apertus model, as an early example. His prescription includes public compute programs funding fully open reference models, neutral foundations holding ground the way they did for internet protocols, commercial participation from companies that benefit from commodity models, and philanthropic funding for the evaluation and audit infrastructure that nobody else will pay for.
He also wants the ecosystem to move past open weights toward fuller transparency. "It's hard to fully trust a model with decisions if you can't tell how it was trained or what it was evaluated against," he said.
For enterprise buyers, the practical takeaway is a pricing test disguised as a capability question. Paying for a closed frontier model currently buys about a four-month head start at roughly five times the cost per task, and that premium is only justified on deadline-sensitive or expert-tier work. Everything else is a candidate for migration to open weights, particularly as the performance gap keeps compressing on an accelerating timeline. Investors and procurement officers alike should treat the 4.4-month figure as a moving target: it has been shrinking, and there is no evidence in this data that the trend is stalling. The bigger structural risk, per Mozilla's own framing, is not the price gap closing but who ends up owning the default infrastructure once it does.
Tags
Original Sources
Exclusive: Paying for frontier AI models buys 4-month head start at 5x the cost
↗ https://arstechnica.com/ai/2026/09/exclusive-open-chinese-models-close-gap-with-silicon-valleys-frontier-ai-models
About the author
Marcus began tracking AI's market implications in 2016, noticing AI-related patent filings accelerating ahead of earnings upgrades before most of the sell-side had caught on. A former fixed-income quantitative analyst, he spent two decades building models that priced risk across emerging markets before pivoting to cover the economic impact of AI full-time. His writing translates opaque technical developments into clear risk/reward terms — and he's rarely diplomatic about the gap between AI valuations and underlying fundamentals. He believes most market participants still underestimate AI's long-run deflationary effect on knowledge work.
More from The Analyst →This Week's Edition
16 September 2026
31 articles
Related Articles
Related Articles
More Stories
© 2026 Cedar & Bloom. All rights reserved.