Google launches Gemini 3.7 Flash: No retraining, just a new inference engine, at half the price
- Google shipped the next Flash model just three weeks later, but it wasn't retrained—only the inference algorithm changed.
- It tops 9 of 20 official benchmarks and costs less than 40% of Claude Sonnet 5's price, but it regressed on one metric compared to its predecessor.
- "Half off" has a date attached. Miss it, and your January bill doubles.
What's new in Gemini 3.7 Flash
Google has released Gemini 3.7 Flash, a workhorse modelA flagship model is like calling in a team of specialists—you pay per consultation. A workhorse model is like hiring a full-time assistant: less powerful per task, but affordable enough to keep running all day. built for coding and agents. The flagship chases benchmark glory; this one handles the volume—cheap, fast, and good enough for tens of thousands of calls a day.
It arrives just three weeks after Gemini 3.6 Flash, at half the price, with broad gains across coding, knowledge work, and web development.
The big change: No retraining, just a new inference algorithm
A three-week iteration cycle is unusually fast. The reason: this model wasn't retrained at all. Gemini 3.7 Flash is built directly on 3.6 Flash, keeping the architecture, training data, hardware, and software identical. The only change specific to 3.7 is an algorithmic improvement in its core inference engine.
The next-gen model is a separate track. When Google launched 3.6 Flash three weeks ago, it mentioned that its most ambitious pretraining run yet was already underway for Gemini 4. 3.7 Flash is an algorithmic iteration on the same base, released before that pretraining run completed.
What got better: Coding, web dev, documents, and workflows
Coding: noticeably higher first-try success rate
Debugging and troubleshooting are better. The rate of writing production-ready code on the first try is up significantly. Production-grade code quality rose from 34.4% to 43.6%, edging out Claude Sonnet 5's 42.7% and GPT-5.6 Terra's 41.3%. For long-horizon software engineering—tasks that require many iterative rounds—the score jumped from 48.6% to 65.3%. Terminal-based work also improved, from 78.0% to 85.8%.
Web development: usable pages from fewer prompts
Generated layouts are more complete and functional in one pass, requiring fewer rounds of revision. When given a screenshot, image, or a full design system as a reference, the output aligns much more closely with the original. The web development Elo rating rose from 1538 to 1588—the highest among the four models compared (Claude Sonnet 5: 1541, GPT-5.6 Terra: 1523).
Complex documents: better comprehension in finance, law, and biology
These fields share dense, lengthy documents layered with jargon. Professional PDF understanding jumped from 22.0% to 34.0%—an increase of over half. Precise information retrieval in long contexts hit 97.0%, also the highest among the four.
Business process automation: nearly doubled
This benchmark feeds a model a real corporate workflow and asks it to complete the steps autonomously. The score went from 17.0% to 30.4%. For context, Claude Sonnet 5 scored 10.7% and GPT-5.6 Terra scored 23.6%—this is where 3.7 Flash beat the competition by the widest margin.
More pleasant to use
Beyond benchmarks, Google highlighted a few quality-of-life improvements: it's better at getting unstuck, asks for clarification when intent is ambiguous, and follows instructions more precisely. It invests more effort in multi-step planning and tool use, which means fewer instances of you needing to supervise or start over.
Four official demo videos
These demos share one thing: none is a simple back-and-forth. The model acts as an orchestrator, managing a swarm of subtasks and driving the job from start to finish.
Full benchmark table: 9 out of 20 wins
The official table includes 20 metrics. Opponents are its predecessor 3.6 Flash, Claude Sonnet 5, GPT-5.6 Terra, and Meta's Muse Spark 1.2. Gemini 3.7 Flash took the top score in 9 of these: production-grade code quality, web development, business process automation, complex legal workflows, professional PDF understanding, long-video understanding, long-context information retrieval, multidisciplinary expert reasoning, and real-world biology research tasks.
You can switch the comparison on this table. By default, it's matched against the previous 3.6 Flash, but click the buttons above to compare against another model. The same 20 rows will be re-scored for wins and losses.
| Metric | 3.7 Flash | Opponent |
|---|---|---|
| Composite intelligence indexArtificial Analysis | 56 | 52555757 |
| Production-grade code qualityFrontierCode 1.1 Main | 43.6% | 34.4%42.7%41.3%N/A |
| Long-horizon software engineeringDeepSWE v1.1 | 65.3% | 48.6%53.8%69.6%54.9% |
| Web developmentCode Arena · Elo | 1588 | 1538154115231535 |
| Terminal-based codingTerminal-bench 2.1 | 85.8% | 78.0%80.4%87.4%82.9% |
| General agent capabilityTerminal-bench 3.0 | 14.9% | 5.4%14.6%20.8%N/A |
| Business process automationAutomationBench | 30.4% | 17.0%10.7%23.6%N/A |
| Knowledge workGDPVal-AA v2 · Elo | 1525 | 1422159815781628 |
| Complex legal workflowsHarvey LAB-AA | 90.7% | 85.1%90.1%85.2%N/A |
| Professional PDF understandingGDP.pdf | 34.0% | 22.0%28.0%24.7%16.0% |
| Complex chart understanding, no toolsCharXiv Reasoning | 84.5% | 85.2%77.0%85.9%N/A |
| Complex chart understanding, with toolsCharXiv Reasoning | 88.7% | 89.4%88.3%N/AN/A |
| Long-video understandingLVBench | 85.4% | 84.2%68.5%78.9%N/A |
| Long-context information retrievalGDM-MRCR v2 · 128k | 97.0% | 91.8%81.5%93.5%N/A |
| Computer useOSWorld-2.0 | 47.9% | 33.8%N/A50.2%N/A |
| Desktop & system tasksAgent's Last Exam | 26.3% | 24.2%33.3%28.0%N/A |
| Multidisciplinary expert reasoningHLE-Verified | 53.6% | 51.2%31.0%51.1%N/A |
| Bioinformatics research · human-solvableBioMysteryBench | 87.1% | 80.6%87.5%83.8%N/A |
| Bioinformatics research · hard for humansBioMysteryBench | 43.5% | 41.2%34.1%49.4%N/A |
| Real-world biology researchLABBench2 | 82.1% | 76.1%80.1%81.2%N/A |
Weak spots: Last in knowledge work, failing grade in general agent capability
It has its shortcomings. In knowledge work, 3.7 Flash scored 1525 Elo, placing it last behind Muse Spark 1.2 (1628), Claude Sonnet 5 (1598), and GPT-5.6 Terra (1578). It's not close. This benchmark measures real white-collar tasks—a different skill set from the coding and workflow automation it's designed for.
It also regressed in one area. Complex chart understanding is split into two settings in the table: without tools it scores 84.5% versus 85.2% for the previous model; with tools it scores 88.7% versus 89.4%. Both dropped by 0.7 points. This is the only metric of the 20 where 3.7 Flash scores lower than 3.6 Flash.
The most telling row is general agent capability. 3.7 Flash's 14.9% is nearly triple the previous 5.4%, which sounds impressive. But on a 0-to-100 scale, the picture is sobering:
Safety assessment: Cyber attacks hit the warning line
Google's frontier safety framework sets two thresholds for each risk domain: a lower "warning" line and an upper "capability red line." Crossing the red line triggers additional deployment restrictions. For 3.7 Flash, the conclusion is: cyber attacks reached the warning line but not the red line. The same applies to chemical, biological, radiological, and nuclear (CBRN) risks: expert red teams could elicit accurate, actionable information across a full attack chain, but because average scores were low and elicitation required explicit expert guidance, it didn't cross the line. Mitigation measures remain in place for both.
There's also a notable conclusion unrelated to capability: 3.7 Flash's situational awareness is stronger than Gemini 3.1 Pro's. It can correctly detect that it is being tested, though it can't circumvent the constraints of the test environment.
Other limitations: it can hallucinate, occasionally runs slow or times out, and its knowledge cutoff is listed as March 2026, with a caveat. It has more recent information in some domains, but in others its knowledge only extends to January 2025. This matters for retrieval and fact-checking—don't just look at the "March" date.
Price cut to half—but only until the end of the year
3.6 Flash launched three weeks ago at $1.50 per 1M input tokens and $7.50 per 1M output tokens. 3.7 Flash now costs $0.75 and $3.75—exactly half.
3.6 Flash has also been reduced to $0.75 and $3.75, bringing both generations to the same price. If you're still on 3.6, switching costs nothing extra.
But this price only lasts until December 31, 2026. On January 1, 2027, both Flash models will revert to $1.50 / $7.50—a double. Cost projections based on the $0.75 rate need to be recalculated at 2× in about four and a half months.
The pricing row on the same table explains the rationale: Claude Sonnet 5 is $2.00 / $10.00, GPT-5.6 Terra is $2.00 / $12.00, and Muse Spark 1.2 is $1.25 / $4.25. At $0.75 / $3.75, 3.7 Flash undercuts all three while scoring just 1 point below the highest composite intelligence index.
The real bill is calculated per task
But the price per million tokens isn't what developers actually end up paying.
An agent completing a task goes through many rounds: reading a file, thinking, calling a tool, observing the result, thinking again, calling again. The total number of tokens burned is determined by the model itself. A verbose model can run up a higher bill despite a lower unit price. Crudely, thinking configurations also raise accuracy while burning more tokens. So what you should care about is the average cost to complete a task, not the sticker price per million tokens. It's like taking a cab: you care about the total fare from home to the office, not the per-kilometer rate. A driver who takes the scenic route isn't cheap, even at a lower rate.
The official chart has exactly this on the horizontal axis: average cost per task, with expensive on the left and cheap on the right. In the top-right corner sits "most efficient." Each model is a line, with dots representing its different thinking configurations. The three Flash generations are connected by a light blue band marching toward the top right: higher scores and lower cost per task, simultaneously.
Reading the band: 3.5 Flash costs roughly $7 per task and gets 37% right. 3.6 Flash costs around $3 and gets about half right. By 3.7 Flash, it's just over $2 per task with a 65% success rate. Meanwhile, the highest-scoring claude-opus-5 on the chart reaches 74%, at a cost of over $10 per task. Even dialed down to its most efficient setting, it takes around $5 to hold onto that 70% range.
How to use it: available today, three entry points
Gemini 3.7 Flash is generally available as of launch day. Each type of user has its own door.
Spark is Google's personal agent announced at I/O. It works for you 24/7; you set the direction, it does the work. As of today, it's powered by 3.7 Flash, with improvements in two areas: more accurate use of Google Workspace tools, and higher-quality outputs for complex, multi-step workflows like merging scattered files, drafting emails, and updating status documents.
The launch page also includes testimonies from 12 customers. The list itself is informative: Box, Databricks, Harvey, Hebbia, LangChain, Browser Use, Cartwheel, Emergent, Nunu.ai, Open Code, Pydantic, and Stanford University's Biology Department. Enterprise document handling, legal, data platforms, agent frameworks, and browser automation—all workloads that involve long, continuous model calls.
In Box's evaluation with real enterprise knowledge work, Gemini 3.7 Flash is more accurate and significantly faster than the previous generation, with the biggest gains on the hardest analytical tasks. Yashodha Bhavnani, VP of Product, Box AI
Gemini 3.7 Flash launches: A new model in 3 weeks, via a new algorithm, not a retrain
Google released its coding/agent model on August 13. Here's the mechanics, the report card, and the year-end pricing cliff, all on one page with visuals.
↓ One page to read it all · Includes one animated chart
On August 13, Google released Gemini 3.7 Flash, a workhorse model for coding and agent workloads—cheap, fast, and built for high-volume calls. It tops 9 of 20 official benchmarks. All figures come from Google's own model card and blog; there's no third-party reproduction yet.
No retraining was done. Architecture, training data, hardware, and software are all carried over from 3.6 Flash. The only change is a new core inference algorithm.
Coding, web development, complex documents, and workflow automation saw the biggest jumps. On the same table, there are also areas where it trails competitors, or even its own predecessor.
3.7 Flash's price is half of 3.6 Flash's launch price, and 3.6 has dropped to match. But this is a promo rate that ends December 31, 2026. On January 1, both models snap back to their original prices.
$1.50 / $7.50
$0.75 / $3.75
The sticker price per million tokens isn't the final bill. An agent uses many rounds to finish a task, so what matters is the average cost to complete one.
is out!
- × No retraining
- × No new architecture
- × No new training data
- × No new hardware/software
core inference
algorithm changed
$0.75/$3.75!
Price snaps back 1/1
and thinks
isn't the real cost
per completed task
Mark the date
