
Share
Google's latest Flash release trades peak intelligence for speed and price discipline, arriving in three configurable tiers that sharpen the tradeoff between capability and cost per task.
Google has released Gemini 3.8 Flash, a three-tier model family that continues the company's strategy of segmenting its Gemini line by reasoning effort rather than by a single monolithic release. The models, tagged high, medium, and low, arrived in September 2026 and are available through Google's own provider infrastructure. This is not a flagship launch. It is a calibration exercise, and the numbers tell that story clearly.
On Artificial Analysis's Intelligence Index, the top configuration, Gemini 3.8 Flash (high), scores 41. The medium variant scores 40, and the low variant comes in at 34. These are respectable but unremarkable figures against a field of 652 models the benchmarking firm tracks. For context, Google's own Gemini 3.1 Pro Preview posts a score of 30 on a different comparison set, while smaller Gemma 4 variants land in the mid-teens. Flash sits in the middle of the pack by design, not by accident.
The more interesting data point is price. All three tiers carry the same list price of $0.60, according to the pricing table on Artificial Analysis. Yet the cost per Intelligence Index task tells a different story once reasoning overhead is factored in. The medium configuration delivers the lowest cost per task at $0.93, while the high configuration costs $1.24 per task, a 1.3x premium for the extra reasoning depth. That gap matters more than the headline list price, because it captures the actual token burn required to complete a standardized task, not just the sticker price per million tokens.
Output speed favors the high-effort configuration, which streams at 345 tokens per second, the fastest of the three tiers based on available data. Time to first answer token also favors the high tier at 30.49 seconds, though that figure reflects a model that is doing more reasoning work before it starts producing visible output, not necessarily a snappier user experience end to end. Speed figures for the medium and low tiers were not published in the release data, an omission that limits how confidently buyers can compare the full lineup on latency.
Gemini 3.8 Flash is a reasoning model, per Artificial Analysis's classification, and a non-reasoning variant may exist alongside it though it is not detailed in this release. The three tiers, high, medium, and low, correspond to different reasoning budgets Google allows the model to spend before answering. This is now a familiar pattern across the industry: rather than training and shipping three separate model weights, providers increasingly expose a single model with a dial that trades latency and cost for accuracy.

The practical implication for buyers is that model selection becomes a tuning exercise rather than a binary choice. A finance team running high-volume, low-stakes classification tasks might default to the low tier at an intelligence score of 34, accepting some quality loss in exchange for speed and lower reasoning overhead. A team running complex multi-step analysis might pay the 1.3x cost premium for the high tier's extra six points of intelligence. Whether that six-point gain justifies the cost difference depends entirely on the task, and Artificial Analysis's benchmarking makes that tradeoff visible rather than theoretical.
All three configurations share the same input and output specifications. Each supports text, image, speech, and video as input modalities, with text-only output. The context window sits at one million tokens, roughly 1,500 pages of standard A4 text, unchanged across all three reasoning tiers. That consistency suggests Google is treating the reasoning dial as an inference-time setting layered on top of a shared underlying model, rather than shipping genuinely distinct architectures.
Positioning against competitors is harder to assess from the release data alone, since the comparison chart references providers including Anthropic, OpenAI, Meta, Alibaba, Z AI, Space, xAI, and IBM without disclosing their specific scores in this excerpt. What is clear is that Artificial Analysis places Gemini 3.8 Flash near the "Pareto line" on its intelligence-versus-cost chart, in the "most attractive quadrant," a designation the firm reserves for models that deliver strong intelligence per dollar rather than the highest absolute score. That is a meaningful signal for cost-conscious buyers, even if it does not settle the question of best-in-class performance.
The Intelligence Index itself draws on ten evaluations, including AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1. Breakdowns by domain, covering finance and accounting, strategy and operations, legal, healthcare, engineering, and economics, are available through Artificial Analysis's capability indices, though granular scores for Gemini 3.8 Flash across those domains were not included in this release summary.
Gemini 3.8 Flash is not chasing the top of the leaderboard. It is a pricing and efficiency play aimed at high-throughput, cost-sensitive workloads, positioned in the attractive quadrant of the intelligence-to-cost curve rather than at its peak. Investors and enterprise buyers evaluating Google's AI stack should watch the medium tier's $0.93 cost per task as the most relevant figure here, since it represents the best efficiency point in the lineup. The bigger question is whether Google can sustain this segmentation strategy as competitors from Anthropic, OpenAI, and Alibaba push their own tiered releases, tightening the margin for differentiation on price alone. Six months of usage data across enterprise deployments will say more than any single benchmark score.
Tags
Original Sources
Gemini 3.8 Flash Models - Intelligence, Performance & Price Comparison | Artificial Analysis
↗ https://artificialanalysis.ai/models/releases/gemini-3-8-flash
About the author
Marcus began tracking AI's market implications in 2016, noticing AI-related patent filings accelerating ahead of earnings upgrades before most of the sell-side had caught on. A former fixed-income quantitative analyst, he spent two decades building models that priced risk across emerging markets before pivoting to cover the economic impact of AI full-time. His writing translates opaque technical developments into clear risk/reward terms — and he's rarely diplomatic about the gap between AI valuations and underlying fundamentals. He believes most market participants still underestimate AI's long-run deflationary effect on knowledge work.
More from The Analyst →This Week's Edition
18 September 2026
25 articles
Related Articles

Medicare's AI Push Meets a Discovery Problem for Health Chatbots
Products & Applications · 5 min

Intermountain Health Closes $795M Deal for Idaho Hospitals, Ending Surgery Partners' Joint Ownership
Products & Applications · 5 min

Microsoft's Internal AI Rollout Offers a Data Point on Enterprise Transformation Economics
Finance & Markets · 5 min
Related Articles

Medicare's AI Push Meets a Discovery Problem for Health Chatbots
Products & Applications · 5 min

Intermountain Health Closes $795M Deal for Idaho Hospitals, Ending Surgery Partners' Joint Ownership
Products & Applications · 5 min

Microsoft's Internal AI Rollout Offers a Data Point on Enterprise Transformation Economics
Finance & Markets · 5 min
More Stories
© 2026 Cedar & Bloom. All rights reserved.