Product Launch · Xiaohu's Take

Google launches Gemini 3.7 Flash: No retraining, just a new inference engine, at half the price

The previous version came out just three weeks ago. The new price is half, but it's a promo that expires at the end of the year.
Quick Take
  • Google shipped the next Flash model just three weeks later, but it wasn't retrained—only the inference algorithm changed.
  • It tops 9 of 20 official benchmarks and costs less than 40% of Claude Sonnet 5's price, but it regressed on one metric compared to its predecessor.
  • "Half off" has a date attached. Miss it, and your January bill doubles.
Benchmarks, pricing, and demos in this article come from Google's own blog and official model card. Results were measured by Google, and competitors were selected and presented by Google. We have not independently reproduced these results.
Launch

What's new in Gemini 3.7 Flash

Google has released Gemini 3.7 Flash, a workhorse modelA flagship model is like calling in a team of specialists—you pay per consultation. A workhorse model is like hiring a full-time assistant: less powerful per task, but affordable enough to keep running all day. built for coding and agents. The flagship chases benchmark glory; this one handles the volume—cheap, fast, and good enough for tens of thousands of calls a day.

It arrives just three weeks after Gemini 3.6 Flash, at half the price, with broad gains across coding, knowledge work, and web development.

3 weeks
Since last version
9 / 20
Official benchmarks won
$0.75 / $3.75
Per 1M input / output tokens
<40%
Of Claude Sonnet 5's price
Inputs
Text, images, audio, video
Context
1M tokens; max 64K output tokens
Config.
Thinking configurationsThink of it as dialing in "how long it deliberates." Longer thinking means better answers, but also higher cost and latency. Shorter thinking is cheaper and faster, but more error-prone. It's like an exam—spending twice as long checking your work does improve your score, but you pay for that extra time.—a single model where you can trade accuracy for speed and cost. Depending on the setting, the cost of a task can vary several-fold.
Mechanics

The big change: No retraining, just a new inference algorithm

A three-week iteration cycle is unusually fast. The reason: this model wasn't retrained at all. Gemini 3.7 Flash is built directly on 3.6 Flash, keeping the architecture, training data, hardware, and software identical. The only change specific to 3.7 is an algorithmic improvement in its core inference engine.

Only this layer is new in 3.7 Flash
Core inference algorithm improvement
Everything else is carried over unchanged
Gemini 3.6 Flash
ArchitectureTraining dataHardwareSoftware
Jul 213.6 Flash launches
3 weeks apart
Aug 133.7 Flash launches
The foundation is carried over entirely; only the inference layer's algorithm changed. That's how a new model ships in three weeks. Redrawn from the official model card.

The next-gen model is a separate track. When Google launched 3.6 Flash three weeks ago, it mentioned that its most ambitious pretraining run yet was already underway for Gemini 4. 3.7 Flash is an algorithmic iteration on the same base, released before that pretraining run completed.

We covered the previous launch in depth
Google released three Gemini models at once, but the long-awaited 3.5 Pro was missing
Improvements

What got better: Coding, web dev, documents, and workflows

Coding: noticeably higher first-try success rate

Debugging and troubleshooting are better. The rate of writing production-ready code on the first try is up significantly. Production-grade code quality rose from 34.4% to 43.6%, edging out Claude Sonnet 5's 42.7% and GPT-5.6 Terra's 41.3%. For long-horizon software engineering—tasks that require many iterative rounds—the score jumped from 48.6% to 65.3%. Terminal-based work also improved, from 78.0% to 85.8%.

Web development: usable pages from fewer prompts

Generated layouts are more complete and functional in one pass, requiring fewer rounds of revision. When given a screenshot, image, or a full design system as a reference, the output aligns much more closely with the original. The web development Elo rating rose from 1538 to 1588—the highest among the four models compared (Claude Sonnet 5: 1541, GPT-5.6 Terra: 1523).

Complex documents: better comprehension in finance, law, and biology

These fields share dense, lengthy documents layered with jargon. Professional PDF understanding jumped from 22.0% to 34.0%—an increase of over half. Precise information retrieval in long contexts hit 97.0%, also the highest among the four.

Business process automation: nearly doubled

This benchmark feeds a model a real corporate workflow and asks it to complete the steps autonomously. The score went from 17.0% to 30.4%. For context, Claude Sonnet 5 scored 10.7% and GPT-5.6 Terra scored 23.6%—this is where 3.7 Flash beat the competition by the widest margin.

Production-grade code qualityFrontierCode 1.1 Main
3.6 Flash34.4%
3.7 Flash43.6%
Long-horizon software engineeringDeepSWE v1.1
3.6 Flash48.6%
3.7 Flash65.3%
Professional PDF understandingGDP.pdf
3.6 Flash22.0%
3.7 Flash34.0%
Business process automationAutomationBench
3.6 Flash17.0%
3.7 Flash30.4%
The full bar represents 100%. These are solid gains, but the absolute numbers are still mid-pack. For tasks where a model must "see a whole job through to the end," no current model scores above a passing grade. Data from the official model card, redrawn by us.

More pleasant to use

Beyond benchmarks, Google highlighted a few quality-of-life improvements: it's better at getting unstuck, asks for clarification when intent is ambiguous, and follows instructions more precisely. It invests more effort in multi-step planning and tool use, which means fewer instances of you needing to supervise or start over.

Demos

Four official demo videos

These demos share one thing: none is a simple back-and-forth. The model acts as an orchestrator, managing a swarm of subtasks and driving the job from start to finish.

From one sentence to a playable 3D game. 3.7 Flash works with Nano Banana to generate characters, items, and textures in real time as you play. Video from Google's launch page.
One-shot generation of a complete, interactive landing page. 3.7 Flash acts as the lead, delegating tasks to sub-agents and handing off the scroll-linked parallax effects to Gemini Omni. Video from Google's launch page.
We've covered both of its partners in the video
Google launches Nano Banana 2 Lite and video model Omni Flash: 4-second image gen, the fastest and cheapest in the series
Training a robot with 3.7 Flash. Three agents form a feedback loop: the model watches the video feed, judges whether an action was correct, and the robot learns faster as a result. Video from Google's launch page.
A static PDF annual report becomes an interactive data webpage. The charts are live, and the model synthesizes findings across pages into a summary. Video from Google's launch page.
Benchmarks

Full benchmark table: 9 out of 20 wins

The official table includes 20 metrics. Opponents are its predecessor 3.6 Flash, Claude Sonnet 5, GPT-5.6 Terra, and Meta's Muse Spark 1.2. Gemini 3.7 Flash took the top score in 9 of these: production-grade code quality, web development, business process automation, complex legal workflows, professional PDF understanding, long-video understanding, long-context information retrieval, multidisciplinary expert reasoning, and real-world biology research tasks.

You can switch the comparison on this table. By default, it's matched against the previous 3.6 Flash, but click the buttons above to compare against another model. The same 20 rows will be re-scored for wins and losses.

Compare Gemini 3.7 Flash to:
vs the previous 3.6 Flash: wins 18, loses 2 out of 20. Both losses are for complex chart understanding, across its two settings. vs Claude Sonnet 5: wins 16, loses 3, with 1 metric not published by the opponent. Losses are in knowledge work, desktop & system tasks, and human-solvable bioinformatics research. vs GPT-5.6 Terra: wins 10, loses 9, with 1 metric unpublished. This is the closest contest, but GPT-5.6 Terra's output price is 3.2× higher than 3.7 Flash's. vs Muse Spark 1.2: wins 4, loses 2. The other 14 metrics were never published by the opponent—the most incomplete data in this table.
Metric3.7 FlashOpponent
Composite intelligence indexArtificial Analysis56 52555757
Production-grade code qualityFrontierCode 1.1 Main43.6% 34.4%42.7%41.3%N/A
Long-horizon software engineeringDeepSWE v1.165.3% 48.6%53.8%69.6%54.9%
Web developmentCode Arena · Elo1588 1538154115231535
Terminal-based codingTerminal-bench 2.185.8% 78.0%80.4%87.4%82.9%
General agent capabilityTerminal-bench 3.014.9% 5.4%14.6%20.8%N/A
Business process automationAutomationBench30.4% 17.0%10.7%23.6%N/A
Knowledge workGDPVal-AA v2 · Elo1525 1422159815781628
Complex legal workflowsHarvey LAB-AA90.7% 85.1%90.1%85.2%N/A
Professional PDF understandingGDP.pdf34.0% 22.0%28.0%24.7%16.0%
Complex chart understanding, no toolsCharXiv Reasoning84.5% 85.2%77.0%85.9%N/A
Complex chart understanding, with toolsCharXiv Reasoning88.7% 89.4%88.3%N/AN/A
Long-video understandingLVBench85.4% 84.2%68.5%78.9%N/A
Long-context information retrievalGDM-MRCR v2 · 128k97.0% 91.8%81.5%93.5%N/A
Computer useOSWorld-2.047.9% 33.8%N/A50.2%N/A
Desktop & system tasksAgent's Last Exam26.3% 24.2%33.3%28.0%N/A
Multidisciplinary expert reasoningHLE-Verified53.6% 51.2%31.0%51.1%N/A
Bioinformatics research · human-solvableBioMysteryBench87.1% 80.6%87.5%83.8%N/A
Bioinformatics research · hard for humansBioMysteryBench43.5% 41.2%34.1%49.4%N/A
Real-world biology researchLABBench282.1% 76.1%80.1%81.2%N/A
The third column is the opponent's score: green with ✓ means 3.7 Flash wins that row, red with ✗ means it loses. Elo is a ranking score—higher is better, not a percentage.
Gemini 3.7 Flash official benchmark comparison table
The official table. The top two rows are pricing; both Gemini entries have an asterisk, explained in the "Price cut to half" section below. Image from Google's official model card.
Two notes on methodology: The web development benchmark appears under two names. The blog post calls it WebDev Arena, while the official table uses Code Arena. The scores are identical (1588 vs 1538). We use the table's naming. Also, for long-horizon software engineering, the blog lists the previous generation's score as 49.0%, while the table and model card both state 48.6%; we use the table's figure.
Weaknesses

Weak spots: Last in knowledge work, failing grade in general agent capability

It has its shortcomings. In knowledge work, 3.7 Flash scored 1525 Elo, placing it last behind Muse Spark 1.2 (1628), Claude Sonnet 5 (1598), and GPT-5.6 Terra (1578). It's not close. This benchmark measures real white-collar tasks—a different skill set from the coding and workflow automation it's designed for.

It also regressed in one area. Complex chart understanding is split into two settings in the table: without tools it scores 84.5% versus 85.2% for the previous model; with tools it scores 88.7% versus 89.4%. Both dropped by 0.7 points. This is the only metric of the 20 where 3.7 Flash scores lower than 3.6 Flash.

The most telling row is general agent capability. 3.7 Flash's 14.9% is nearly triple the previous 5.4%, which sounds impressive. But on a 0-to-100 scale, the picture is sobering:

GPT-5.6 Terra20.8%
Gemini 3.7 Flash14.9%
Claude Sonnet 514.6%
Gemini 3.6 Flash5.4%
The full bar represents 100%. All four models are clustered in the leftmost sliver; the best performer, GPT-5.6 Terra, fills only a fifth. This metric measures "handing a computer to the model and letting it get the job done on its own.” No one is passing yet. Data from the official model card, redrawn by us.
Safety

Safety assessment: Cyber attacks hit the warning line

Google's frontier safety framework sets two thresholds for each risk domain: a lower "warning" line and an upper "capability red line." Crossing the red line triggers additional deployment restrictions. For 3.7 Flash, the conclusion is: cyber attacks reached the warning line but not the red line. The same applies to chemical, biological, radiological, and nuclear (CBRN) risks: expert red teams could elicit accurate, actionable information across a full attack chain, but because average scores were low and elicitation required explicit expert guidance, it didn't cross the line. Mitigation measures remain in place for both.

Capability red line
Crossing this would trigger extra deployment limits. 3.7 Flash doesn't reach it in any domain.
Warning line
Cyber attacks and CBRN both touch this line; mitigation measures will continue.
Below the line
Malicious manipulation and ML R&D acceleration/automation are below this threshold.
Redrawn from the official frontier safety assessment in the model card.

There's also a notable conclusion unrelated to capability: 3.7 Flash's situational awareness is stronger than Gemini 3.1 Pro's. It can correctly detect that it is being tested, though it can't circumvent the constraints of the test environment.

Other limitations: it can hallucinate, occasionally runs slow or times out, and its knowledge cutoff is listed as March 2026, with a caveat. It has more recent information in some domains, but in others its knowledge only extends to January 2025. This matters for retrieval and fact-checking—don't just look at the "March" date.

Pricing

Price cut to half—but only until the end of the year

3.6 Flash launched three weeks ago at $1.50 per 1M input tokens and $7.50 per 1M output tokens. 3.7 Flash now costs $0.75 and $3.75—exactly half.

3.6 Flash has also been reduced to $0.75 and $3.75, bringing both generations to the same price. If you're still on 3.6, switching costs nothing extra.

But this price only lasts until December 31, 2026. On January 1, 2027, both Flash models will revert to $1.50 / $7.50—a double. Cost projections based on the $0.75 rate need to be recalculated at 2× in about four and a half months.

Jul 21
3.6 Flash launches at $1.50 / $7.50
$1.50 / $7.50
Aug 13
3.7 Flash launches; 3.6 drops to match
$0.75 / $3.75
Dec 31
Promo pricing ends for both Flash models
$0.75 / $3.75
Jan 1
Price returns to standard, doubling
$1.50 / $7.50
Bar length represents the price per 1M tokens, input / output, in USD. Data from the official model card footnote and the previous launch blog post.

The pricing row on the same table explains the rationale: Claude Sonnet 5 is $2.00 / $10.00, GPT-5.6 Terra is $2.00 / $12.00, and Muse Spark 1.2 is $1.25 / $4.25. At $0.75 / $3.75, 3.7 Flash undercuts all three while scoring just 1 point below the highest composite intelligence index.

The Real Cost

The real bill is calculated per task

But the price per million tokens isn't what developers actually end up paying.

An agent completing a task goes through many rounds: reading a file, thinking, calling a tool, observing the result, thinking again, calling again. The total number of tokens burned is determined by the model itself. A verbose model can run up a higher bill despite a lower unit price. Crudely, thinking configurations also raise accuracy while burning more tokens. So what you should care about is the average cost to complete a task, not the sticker price per million tokens. It's like taking a cab: you care about the total fare from home to the office, not the per-kilometer rate. A driver who takes the scenic route isn't cheap, even at a lower rate.

The official chart has exactly this on the horizontal axis: average cost per task, with expensive on the left and cheap on the right. In the top-right corner sits "most efficient." Each model is a line, with dots representing its different thinking configurations. The three Flash generations are connected by a light blue band marching toward the top right: higher scores and lower cost per task, simultaneously.

Accuracy ← Higher cost per task Lower cost → Average cost per task claude-opus-5 · ~74% Cost: ~$10+ per task 3.5 Flash · ~37% 3.6 Flash · ~49% 3.7 Flash · ~65%
A simplified version of the official cost curve, showing only the three Flash generations and one high-end reference. Dollar figures were read by eye from the original chart, so the text uses approximations. The original is below.

Reading the band: 3.5 Flash costs roughly $7 per task and gets 37% right. 3.6 Flash costs around $3 and gets about half right. By 3.7 Flash, it's just over $2 per task with a 65% success rate. Meanwhile, the highest-scoring claude-opus-5 on the chart reaches 74%, at a cost of over $10 per task. Even dialed down to its most efficient setting, it takes around $5 to hold onto that 70% range.

DeepSWE V1.1 cost per task compared to accuracy chart
The official chart. The horizontal axis runs from $10 on the left to $0 on the right; cheaper is further right. Each line is a model at its different thinking settings. Gemini 3.7 Flash is highlighted in the top right. Image from Google's launch page, data from Datacurve AI.
Getting Started

How to use it: available today, three entry points

Gemini 3.7 Flash is generally available as of launch day. Each type of user has its own door.

Developers
Run agent-first workflows in Google Antigravity; or access the Gemini API directly from Google AI Studio or Android Studio.
Enterprise
Gemini Enterprise Agent Platform, plus the Gemini Enterprise application.
Consumers
Spark in the Gemini app, which requires a Google AI Pro or Ultra subscription. Available in 160+ countries, excluding the European Economic Area, Nigeria, Switzerland, and the UK.

Spark is Google's personal agent announced at I/O. It works for you 24/7; you set the direction, it does the work. As of today, it's powered by 3.7 Flash, with improvements in two areas: more accurate use of Google Workspace tools, and higher-quality outputs for complex, multi-step workflows like merging scattered files, drafting emails, and updating status documents.

Spark working in a manufacturing scenario. Merging files, drafting emails, and updating status documents without step-by-step instructions. Video from Google's launch page.

The launch page also includes testimonies from 12 customers. The list itself is informative: Box, Databricks, Harvey, Hebbia, LangChain, Browser Use, Cartwheel, Emergent, Nunu.ai, Open Code, Pydantic, and Stanford University's Biology Department. Enterprise document handling, legal, data platforms, agent frameworks, and browser automation—all workloads that involve long, continuous model calls.

In Box's evaluation with real enterprise knowledge work, Gemini 3.7 Flash is more accurate and significantly faster than the previous generation, with the biggest gains on the hardest analytical tasks. Yashodha Bhavnani, VP of Product, Box AI
🧰 Quick Start · Gemini 3.7 Flash
Price$0.75 per 1M input tokens, $3.75 per 1M output tokens; promotional price, expires December 31, 2026, then reverts to $1.50 / $7.50
RequirementsA Google account is enough to try it directly in AI Studio; for volume, use the Gemini API; for agent workflows, use Google Antigravity; for enterprise, use Gemini Enterprise
Source
Introducing Gemini 3.7 FlashGoogle·blog.google·2026-08-13
Our Notes
Benchmarks and pricing are taken from the official model card's comparison table. The web development metric is called "WebDev Arena" in the blog and "Code Arena" in the table; we follow the table. The blog lists the previous generation's long-horizon software engineering score as 49.0%, while the table and model card state 48.6%; we use the table's figure. Dollar amounts on the cost curve were read by eye from the original chart and are presented as approximations. The base composition diagram, four improvement comparison bars, general agent capability bars, price timeline, three-generation Flash positioning chart, and safety tier diagram are redrawn by us from official data. The original price of 3.6 Flash ($1.50 / $7.50) comes from the previous launch blog post. The Box testimony was an image on the launch page; we transcribed it from a screenshot. Attribution for Muse Spark 1.2 comes from our previous coverage of the Muse Spark 1.1 launch.