Skip to main content
Product2026-09-027 min read

Gemini 3.8 Flash Is Now on Multos — The Most Capable Flash Model Google Has Ever Shipped

73.7% on DeepSWE v1.1. A three-point Intelligence Index jump over 3.7 Flash. 304 tokens per second. And it costs exactly the same as the model it just replaced. Gemini 3.8 Flash is live on Multos Starter today.

O
Oghenekaro Oboido

Google shipped Gemini 3.8 Flash on September 2, 2026 — its third Flash release in 43 days — and the numbers landed differently than the usual incremental update. DeepSWE v1.1 went from 65.3% to 73.7%, putting it within 0.3 points of Claude Opus 5 at one-seventh of the cost. Vals Finance Agent and Harvey's Legal Agent both went to 3.8 Flash. Artificial Analysis scored it 59 on its Intelligence Index, up from 56 for 3.7 Flash, and measured throughput at 304.6 tokens per second. The price stayed at $0.75 input and $3.75 output per million tokens. Every spec stayed the same. Only the results changed, and they changed on every published row.

Why This Release Is Different

Most model updates move one benchmark and hold everything else flat. Gemini 3.8 Flash moved all of them. Google's own evaluation table has 14 rows comparing 3.8 Flash against 3.7 Flash, Claude Opus 5, and GPT-5.6 Sol. Gemini 3.8 Flash wins 8 of those 14. Every single row improved over 3.7 Flash. The headline is DeepSWE v1.1 at 73.7% versus 74.0% for Opus 5 — a statistical tie between a model that costs $0.75 per million tokens and one that costs $5.00. The most dramatic individual gain is BioMysteryBench Human Difficult, which went from 43.5% to 56.5%. Terminal-bench 4.0 went from 11.2% to 19.1%. Vals Finance Agent v2, which scores AI on real financial analyst tasks, went to 61.4% — ahead of both Claude Opus 5 (58.6%) and GPT-5.6 Sol (53.8%). Harvey's Legal Agent, which tests complex legal workflows, went to 10.0% versus 8.2% for 3.7 Flash. For professional knowledge work — the kind of multi-step reasoning that most enterprise AI actually does — 3.8 Flash is now the strongest cheap model you can run.

The Benchmark Row Google Won Against Opus 5

There is a number circulating in the launch coverage that is wrong. Multiple write-ups reported Gemini 3.8 Flash at 71.0% on DeepSWE v1.1. Google's evaluation PDF says 73.7%. The difference matters more than it looks: at 71.0% the story is 'Opus 5 comfortably ahead.' At 73.7% it is a statistical tie between a $0.75 model and a $5.00 one. Read anything quoting 71.0% knowing that. The honest picture is that Opus 5 still wins where it matters most for open-ended agent work. Terminal-bench 4.0 goes 51.8% to 19.1% — a gap so wide that even GPT-5.6 Sol finishes ahead of Gemini there. OSWorld-2.0 goes 75.4% to 59.0%. GDPVal-AA v2 Elo is 1824 versus 1545. If you are running an agent that needs to operate a computer freely or navigate general open-ended tasks without supervision, Opus 5 is still the right choice and 3.8 Flash does not change that. What 3.8 Flash does change is bounded professional work: software engineering measured by a harness, finance tasks with defined steps, legal workflows with known structure. For those jobs it now costs one-seventh as much as the best alternative and trades within a point of it.

Speed and What It Means for Multos Agents

Artificial Analysis measured 3.8 Flash at 304.6 output tokens per second in independent testing, compared to 279.4 for 3.7 Flash. Both numbers are fast. The 25 token per second improvement matters less for a single chat exchange and more for the long streaming responses that Multos agentic runs produce. When an agent is reading a codebase, writing a plan, calling tools, and writing back results — all in one turn — the difference between 280 and 305 tokens per second adds up across a session. Time to first token is reported at 13 seconds at the high reasoning setting, which is consistent with other frontier reasoning models. At medium (the default on Multos), response start time is faster. The 1M token context window is unchanged, which means Multos agents can maintain the same long-running sessions with 3.8 Flash that they do with 3.7 Flash. Nothing regresses. The model just produces better output at the same speed profile you already have.

The Pricing Footnote Almost Everyone Missed

The $0.75 input and $3.75 output rates are introductory pricing, and they expire on December 31, 2026. From January 1, 2027 the rates become $1.50 input and $7.50 output. Context caching doubles at the same time from $0.075 to $0.15 per million tokens, and cache storage goes from $0.50 to $1.00 per million tokens per hour. Google states this clearly in a footnote of its own benchmark table, so it is not hidden, but almost every launch article reported the promotional numbers as though they were permanent. If you are modelling costs for a production workload that runs into next year, plan for the January rates. At $1.50 and $7.50, Gemini 3.8 Flash is still cheaper than Claude Opus 5 at $5.00 and $25.00. The value story survives the increase. But the math changes, and you should know that now rather than in January.

Full Spec Sheet

Model ID: gemini-3.8-flash. Released: September 2, 2026. Context window: 1,048,576 tokens (1M). Output limit: 65,536 tokens (64k). Knowledge cutoff: March 2026 (some domains only to January 2025 per the model card). Input modalities: text, image, video, audio, PDF. Output: text only. Thinking levels: low, medium (default), high. Not supported: audio generation, image generation, Live API. Input price: $0.75 per 1M tokens introductory, rising to $1.50 on January 1 2027. Output price: $3.75 per 1M tokens introductory, rising to $7.50. Context caching: $0.075 per 1M tokens plus $0.50 per 1M tokens per hour storage. Artificial Analysis Intelligence Index: 59 (high reasoning), 57 (medium). Throughput: 304.6 output tokens per second.

Why Gemini 3.8 Flash on Multos Is the Obvious Default

The case for using 3.8 Flash on Multos instead of 3.7 Flash is simple: every benchmark improved, the price is the same, and the spec is identical. There is no trade-off to evaluate. You are not giving anything up. You are not paying more. You are getting a model that scores 73.7% on long-horizon software engineering instead of 65.3%, and 61.4% on finance agent tasks instead of 59.0%. The reason Multos offers both models simultaneously is to avoid breaking existing workflows. If you have a prompt or an agent configuration that you have tuned specifically on 3.7 Flash behavior, it will continue to work. But if you are starting a new project, or if you are choosing a default model for your team on Multos, 3.8 Flash is the right answer. It is faster, it is smarter on every published measure, and it costs exactly the same as the model you would otherwise use. Select it from the Google group in the model dropdown. It is available on every Multos plan from Starter upward.

Ready to Get Started?

Gemini 3.8 Flash is live on Multos Starter and above — 73.7% on DeepSWE v1.1, 304 tokens per second, same $0.75 input price as 3.7 Flash. Open the model dropdown, find it under Google, and run your next project on the most capable Flash model Google has ever shipped.

Start Building Free