Google released Gemini 3.8 Flash on September 2, 2026, three weeks after Gemini 3.7 Flash and one day after Anthropic shipped Claude Fable 5.1. The pitch is blunt: frontier-class results on coding, finance, and legal agent benchmarks at a fraction of rivals’ prices, with promotional pricing of $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026.
The numbers behind that pitch are more interesting than the pitch itself. Gemini 3.8 Flash beats Claude Opus 5 on several of Google’s chosen benchmarks and trails it badly on others, including the hardest current coding test. Here is what changed, what the model is actually good at, what it costs, where to use it, and why Google is shipping a new Flash model every three weeks.
What Google Released
Gemini 3.8 Flash is the next step in the Gemini 3 family, described in Google’s official model card as delivering advances across software engineering and agentic knowledge workflows, built for cost-effective scaling of production agents. It is available immediately in the Gemini app for AI Pro and Ultra subscribers, in AI Mode, in Gemini for Google Sheets, in the Antigravity coding environment, and to developers through Google AI Studio, the Gemini API, and the Gemini Enterprise Agent Platform.
Specifications at a glance
| Spec | Gemini 3.8 Flash |
|---|---|
| Context window | Up to 1 million tokens |
| Maximum output | 64,000 tokens |
| Inputs | Text, images, audio, video |
| Effort levels | Customizable, trading quality against cost and latency |
| Knowledge cutoff | March 2026 for some domains, January 2025 for others |
| Pricing (promo through Dec 31, 2026) | $0.75 input, $3.75 output per million tokens |
Reporting on the launch puts regular pricing after the promotion at double those rates, $1.50 and $7.50, though Google’s own materials lead with the introductory numbers.
What Is New vs Gemini 3.7 Flash
It works harder, and spends more tokens doing it
Google has not disclosed size or architecture changes; the difference it describes is behavioral. It says 3.8 Flash works harder on complex tasks: it runs additional reasoning steps and calls tools iteratively, sometimes spending more tokens to reach a better answer. That is the same direction Anthropic took with Fable 5.1 a day earlier, and the direction OpenAI has described for its still-unreleased Astra: models that persist through long, multi-step jobs rather than answering once and stopping. The trade-off is built in: a cheap model that thinks longer is not always a cheap task, which is why effort levels matter more with each release.
Safety: assessed by inheritance, with one regression
The model card reports similar safety and tone to 3.7 Flash after manual red teaming, with one exception it states plainly: automated safety performance across non-English languages regressed by 5.4 percentage points relative to 3.7 Flash. On frontier risks, Google did not run a full new assessment for 3.8 Flash. It evaluated 3.7 Flash against its Frontier Safety Framework, found no tracked or critical capability levels reached, and concluded that because 3.8 Flash shows no meaningful new capabilities in those domains, it is also unlikely to cross them. That is a reasonable inference, but an inference rather than a fresh evaluation.
Gemini vs Claude Opus 5: Where 3.8 Flash Wins
Google’s launch table compares 3.8 Flash with Claude Opus 5 and, on some tests, GPT-5.6 Sol. All figures below are Google’s own runs, reported by Google; independent replication is still ahead. On the professional and analytical benchmarks, the cheap model leads.
| Benchmark | Gemini 3.8 Flash | Claude Opus 5 | GPT-5.6 Sol |
|---|---|---|---|
| Vals Finance Agent v2 | 61.4% | 58.6% | 53.8% |
| Harvey Legal Agent | 10.0% | 6.7% | 2.5% |
| Terminal-Bench 2.1 (coding) | 89.4% | 89.1% | n/a |
| CharXiv Reasoning (charts) | 86.2% | 83.7% | n/a |
| LVBench (long video, agentic mode) | 87.8% | 75.4% | n/a |
| Humanity’s Last Exam, verified | 54.9% | 54.4% | n/a |
| LABBench2 (lab science) | 86.2% | 84.2% | n/a |
| BioMysteryBench, human difficult | 56.5% | 49.4% | 44.7% |
The finance and legal results are the ones enterprises will notice: on Vals Finance Agent, 3.8 Flash edges a model that costs roughly six to seven times more per token. For high-stakes domains like these, a vendor benchmark is a reason to test, not a reason to trust; expert review and your own evaluations still apply. The video result uses the agentic video mode Google shipped the day before. And the widely shared 89.4 percent figure is Terminal-Bench 2.1, an older version of the agentic coding benchmark, which matters for the next section.
Gemini vs Claude: Where 3.8 Flash Loses
Google published the losses too, which deserves credit. Some are narrow, such as DeepSWE, where 3.8 Flash lands within a third of a point of Opus 5; others are large.
| Benchmark | Gemini 3.8 Flash | Gemini 3.7 Flash | Claude Opus 5 |
|---|---|---|---|
| Terminal-Bench 4.0 (current coding) | 19.1% | 11.2% | 51.8% |
| OSWorld 2.0 (computer use) | 59.0% | 50.6% | 75.4% |
| DeepSWE v1.1 (software engineering) | 73.7% | 65.3% | 74.0% |
| GDPval-AA v2 (knowledge work) | 1545 | 1482 | 1824 |
| BioMysteryBench, human solvable | 88.8% | 87.1% | 90.1% |
The Terminal-Bench gap is the story inside the story. On version 2.1, 3.8 Flash matches Opus 5; on version 4.0, the current and much harder edition, it scores 19.1 percent against 51.8 for Opus 5 in Google’s run. For context only, Anthropic’s own runs put Opus 5 at 52.3 and Fable 5.1 at 55.8, but a different evaluator with different settings and safeguards is not directly comparable. Same benchmark family, opposite conclusions, depending on which version a chart chooses to show. The honest reading: on Google’s evidence, 3.8 Flash is a strong analyst and a capable coder on well-trodden tasks, and not yet a frontier agentic coder on the hardest work.
Gemini 3.8 Flash Pricing vs Opus 5 and GPT-5.6 Sol
| Model | Input per million tokens | Output per million tokens |
|---|---|---|
| Gemini 3.8 Flash (promo through 2026) | $0.75 | $3.75 |
| GPT-5.6 Sol (promo through late Nov) | $4 | $20 |
| Claude Opus 5 | $5 | $25 |
| Claude Fable 5.1 | $10 | $50 |
Per token, Gemini 3.8 Flash is roughly a fifth of Sol’s price and a sixth or seventh of Opus 5’s. Two caveats: the promotional rate expires at year end, and a model that works harder uses more tokens per task, so cost per finished job is what matters, and Google has not published per-task cost curves the way Anthropic did for Fable 5.1. For high-volume, well-defined workloads, the arithmetic still favors Flash by a wide margin.
Why Google Ships a Flash Model Every Three Weeks
Gemini 3.6 Flash, 3.7 Flash on August 13, 3.8 Flash on September 2. According to reporting sourced to the Wall Street Journal, the cadence is a strategy born of a weakness: Google has lagged Anthropic and OpenAI in AI-assisted coding, and the lighter Flash architecture lets multiple research teams iterate with reinforcement learning at once, while the larger Pro line has suffered scrapped candidates and slipped timelines. The same reporting says engineers on Google’s internal developer platform preferred the new model over Anthropic’s Opus in testing, and that new DeepMind leadership has pushed for a faster pace of execution. Read that way, 3.8 Flash is less a product launch than a sprint milestone.
The Trusted-Tester Cyber Variant
Alongside the public model, Google announced Gemini 3.8 Flash Cyber, a variant built for autonomous vulnerability discovery and automated patching, offered through what Google calls the Fairwind Program to trusted government authorities, critical infrastructure operators, and software maintainers. Google reports frontier-level results on the CyberGym vulnerability benchmark, a 47.2 percent pass rate on CWE-Bench, and improved resistance to prompt injection. That mirrors the pattern set this week by Anthropic’s Mythos 5.1 and by OpenAI’s handling of Astra: the most capable security behaviors ship to vetted defenders first, and the general public gets the safeguarded version.
The Same-Day Context: A Price War at the Frontier
Gemini 3.8 Flash did not land alone. Alibaba shipped an upgraded Qwen 3.8 Max the same day, Fable 5.1 arrived the day before with a 75 percent cut to cache-read pricing, and OpenAI’s Astra remains unreleased but publicly confirmed. The pattern is identical across labs: capability rises while the price of good-enough intelligence falls, and the competition is shifting from smartest model to cheapest acceptable result.
How to Choose Between Flash, Fable, and Sol
With three major models shipping in a week, a simple rule covers most cases. Use Gemini 3.8 Flash when the work is high-volume or analytical: document and spreadsheet analysis, chart reading, long-video questions, finance and legal research, routine code changes, and anything where cost per call dominates. Use Claude Fable 5.1 or Opus 5 when the task is a long, autonomous engineering job or a multi-step agent workflow where the Terminal-Bench 4.0 and OSWorld gaps translate into real failures. Use GPT-5.6 Sol where its promotional pricing and ecosystem already fit your stack. On Google’s own benchmarks, the cheap model handles many routine and analytical tasks well and falls short on the hardest agentic work, and that is a reasonable way to deploy it until independent results arrive.
What It Means for Everyday Users
For Gemini app subscribers, the change is invisible and immediate: the model is live in the app at no extra cost, with its strongest gains on analytical questions and chart reading. The long-video result was measured in agentic video mode, which Google has so far promised for the Gemini app only as coming soon, so do not expect that particular gain in the app yet. The Sheets integration is the sleeper feature for office workers, since 3.8 Flash’s strongest results are precisely in the finance and data analysis tasks spreadsheets are used for, and Antigravity users get it as a coding assistant for routine tasks. Two practical takeaways follow. Effort levels are now a standard dial across Google, Anthropic, and OpenAI, and learning when to turn them down saves real money. And benchmark charts are marketing until you check which version of the test they cite, as the Terminal-Bench split shows.
The Skill That Compounds Across Every Release
A new frontier-class model every week is the new normal, and the people who benefit are those who can quickly judge what a model is good for, brief it precisely, and check its output. Those habits transfer from Flash to Fable to whatever ships next Wednesday. Coursiv builds them with step-by-step guides, short daily lessons, and hands-on practice with AI tools, designed for busy people without a technical background. Check the official site for current course details and pricing.
What to Watch Next
Watch for independent Terminal-Bench 4.0 and DeepSWE results, which will show whether the coding gap is as narrow as DeepSWE suggests or as wide as Terminal-Bench 4.0 does. Watch the price after December 31, when the promotional rate is due to double. Watch whether the three-week cadence holds, and whether a Gemini Pro release finally arrives to compete at the top of the table rather than the bottom of the price list.
The Bottom Line
Gemini 3.8 Flash is the strongest argument yet that frontier-adjacent intelligence is becoming cheap, and the clearest example yet that benchmark tables need reading with the version numbers on. It wins where analysis matters and loses where the hardest agentic coding lives. The price looks highly competitive; whether the cost per finished task and the quality hold up is for independent testing to show.