
Share
Google's latest small model splits into high, medium, and low reasoning variants, trading a few intelligence points for sharply lower costs. Here's what the benchmarks actually show, and who should care.
Google quietly rolled out Gemini 3.8 Flash in September 2026, and the interesting part isn't the headline model, it's the fact that there are three of them. Gemini 3.8 Flash (high), (medium), and (low) all ship as reasoning variants of the same underlying model, each tuned for a different point on the intelligence-versus-cost curve. If you've been treating "reasoning effort" as a runtime toggle rather than a separate deployment decision, this release is a good reminder that it's becoming a first-class product axis for Google.
Here's the breakdown, per Artificial Analysis benchmarking:
That one-point gap between high and medium is almost a rounding error on a 41-point scale. What's not a rounding error is the pricing. Medium runs $0.93 per Intelligence Index task, the cheapest of the three, while high costs $1.24. Artificial Analysis notes prices vary up to 1.3x across the tier, which lines up with that math. Listed API pricing has all three variants sharing a $0.6 base rate, so the real cost difference shows up in how many reasoning tokens each tier burns through to get to an answer, not in the sticker price per token.
Latency tells a similar story. Gemini 3.8 Flash (high) posts a 16.38 second time-to-first-answer-token, the lowest reported in the group. For anything interactive, chat interfaces, agentic loops with tight turnaround expectations, that number matters more than the raw intelligence score.
The Intelligence Index itself is worth a quick refresher if you haven't been tracking Artificial Analysis's methodology. Version 4.3.2 blends ten evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1. It's a composite score meant to capture reasoning, coding, and real-world task competence rather than any single skill, and Gemini 3.8 Flash (high) currently ranks 28th out of 673 tracked models on it.
That ranking is the context that matters here. A score of 41 doesn't put Gemini 3.8 Flash near the frontier, it puts it in the crowded middle tier where cost efficiency, not raw capability, is the differentiator. Artificial Analysis's cost-per-task charts back this up: on the Intelligence Index versus cost scatter plot, Gemini 3.8 Flash sits inside what the firm calls the "most attractive quadrant", the zone where you get reasonable intelligence without paying a premium for it. It's positioned close to the Pareto frontier, the line representing the best available tradeoffs across all tracked models from Google, Anthropic, OpenAI, Meta, xAI, Alibaba, and others.

On capability, the model supports text, image, speech, and video as input modalities, with text-only output. Context window is 1 million tokens, roughly 1,500 A4 pages at 12-point Arial, which is plenty of room for long documents, extended agent transcripts, or multi-file codebases without chunking. It also sits explicitly as a reasoning model in Artificial Analysis's classification, meaning it does internal chain-of-thought style deliberation before producing an answer, and the source notes a non-reasoning variant may exist separately, though it's not part of this comparison.
Where this actually plays out is in domain-specific benchmarking. Artificial Analysis tracks Gemini 3.8 Flash across six capability indices: Finance & Accounting, Strategy & Ops, Legal, Healthcare & Medical, Engineering, and Economics. Each of these blends a subset of the same underlying evaluations (Omniscience, GDPval, Briefcase, HLE, and others) weighted differently for the domain. The exact scores per index weren't broken out in the release data, but the framing is useful: this isn't a model being pitched as a generalist frontier competitor, it's being positioned for teams that need "good enough" reasoning at high volume across specific business functions.
Compare that positioning to siblings in the same family tree. Gemini 3.1 Pro Preview tops out around 30 on the same index scale but presumably targets a different cost bracket. Gemini 3.5 Flash-Lite sits at 22. The Gemma 4 open-weight line (31B, 26B A4B, 12B variants) ranges from 14 to 19. Flash 3.8 clearly slots in as Google's mid-tier workhorse, above the lightweight Gemma releases but well below anything branded "Pro" or aimed at frontier benchmarks.
The three-tier reasoning split is the real story here, not the intelligence score itself. Google is explicitly letting developers trade a single intelligence point (41 versus 40) for a meaningfully lower cost per task (medium's $0.93 versus high's $1.24), which is a genuinely useful lever if you're running high-volume inference where marginal cost compounds fast.
For latency-sensitive applications, high is still your best bet at 16.38 seconds to first token and 293 tokens per second throughput, numbers that outpace whatever medium and low are doing, though Google/Artificial Analysis hasn't published comparable speed figures for those two yet. If you're building anything where users are staring at a loading spinner, that gap is the one to watch, not the intelligence delta.
Ranked 28th of 673 models tracked, Gemini 3.8 Flash isn't chasing the top of the leaderboard. It's chasing the sweet spot on the cost-intelligence curve, and based on where it lands in Artificial Analysis's "most attractive quadrant" plot, it's doing a reasonable job of it. Worth testing against your specific workload before committing, especially if your use case leans heavily on one of those domain capability indices rather than general reasoning.
Tags
Original Sources
Gemini 3.8 Flash Models - Intelligence, Performance & Price Comparison | Artificial Analysis
↗ https://artificialanalysis.ai/models/releases/gemini-3-8-flash
About the author
Kai built ML infrastructure at a Bay Area startup before developing an obsession with transformer architectures and inference optimisation that eventually pulled him out of product work entirely. A stint at a compute research lab sharpened his instinct for what actually matters in a model release versus what is marketing. He writes from the inside — from the perspective of someone who has debugged the systems he is describing at three in the morning. He is allergic to hype and instinctively drawn to the unglamorous plumbing questions that everyone else skips over.
More from The Engineer →This Week's Edition
25 September 2026
31 articles
Related Articles

China Telecom AI Releases Xing4.0-29B-A4B, a 29B-Parameter Agentic Model That Fits on One Consumer GPU
Models & Research · 5 min

Inside the TPC Hackathon: How Researchers Are Building Agentic AI for Supercomputing
Models & Research · 5 min

Meta's Privacy Pivot: Why Muse Needs a Trust Architecture, Not Just a Chatbot
Models & Research · 5 min
Related Articles

China Telecom AI Releases Xing4.0-29B-A4B, a 29B-Parameter Agentic Model That Fits on One Consumer GPU
Models & Research · 5 min

Inside the TPC Hackathon: How Researchers Are Building Agentic AI for Supercomputing
Models & Research · 5 min

Meta's Privacy Pivot: Why Muse Needs a Trust Architecture, Not Just a Chatbot
Models & Research · 5 min
More Stories
© 2026 Cedar & Bloom. All rights reserved.