by datastudy.nl

Tuesday, August 18, 2026

AI

Gemini 3.7 Flash lifts coding scores at half the price

Gemini 3.7 Flash is Google's coding and agent workhorse. It beats 3.6 Flash across every benchmark Google published and undercuts Claude Sonnet 5 on output token price by roughly two-thirds, at $0.75 per million input tokens.

Slope chart comparing Gemini 3.6 Flash to 3.7 Flash across four benchmarks. FrontierCode 1.1 Main rises from 34.4% to 43.6%, DeepSWE v1.1 from 49.0% to 65.3%, AutomationBench from 17.0% to 30.4%, and Terminal-bench 2.1 from 78.0% to 85.8%.
Gemini 3.6 Flash to 3.7 Flash benchmark improvements across coding and agent evaluations. Source: Google DeepMind. Data Today benchmark.

Three weeks. That is the entire shelf life of Gemini 3.6 Flash as Google's workhorse model. On August 13, Google shipped Gemini 3.7 Flash, its latest model tuned for coding and agent workflows, with benchmark jumps large enough to reset expectations for what a mid-tier model can do. The introductory price is $0.75 per million input tokens and $3.75 per million output tokens, half the original cost of 3.6 Flash. That price doubles on January 1, 2027, and anyone building cost-sensitive agent pipelines needs to plan for the switch now, not in December.

Gemini 3.7 Flash is Google's coding and agent workhorse. It beats its predecessor on every benchmark Google published and undercuts Claude Sonnet 5 and GPT-5.6 Terra on output token price by a wide margin. Whether the performance gains translate to your codebase is a separate question, and the introductory pricing expires before the year ends.

What did Google actually ship on August 13?

Gemini 3.7 Flash is a direct successor to 3.6 Flash, built on the same Gemini 3 model family with what Google calls algorithmic improvements to its core reasoning foundation. The model card confirms it supports customizable thinking configurations to control the mix of quality, cost, and latency, a 1 million token context window, and 64K token output. It accepts text, images, audio, and video as input.

The release came three weeks after 3.6 Flash, which itself was a recent model. The cadence tells you something about the competitive pressure Google feels. Anthropic, OpenAI, and Alibaba are all shipping model updates on monthly cycles, and Google is matching that pace with its Flash line. The Flash models are the volume tier: cheaper, faster, and positioned as the model you run in production rather than the one you brag about on leaderboards.

Gemini Spark, Google's personal agent for AI Pro and Ultra subscribers in over 160 countries, switched to 3.7 Flash on launch day. That means Google is dogfooding this model in a consumer-facing agent product from day one, which is a signal about how confident they are in its reliability for multi-step tool use.

How big are the benchmark jumps over 3.6 Flash?

The numbers are the story, and they are large enough to take seriously. On FrontierCode 1.1 Main, a benchmark for production code quality, 3.7 Flash scores 43.6% compared to 34.4% for 3.6 Flash, a 9.2 percentage point jump. On DeepSWE v1.1, which tests long-horizon software engineering, the gap is wider: 65.3% versus 49.0%, a 16.3 point improvement.

The agentic benchmarks show the same pattern. AutomationBench, which measures real-world business workflow automation, jumps from 17.0% to 30.4%. Terminal-bench 2.1, which tests agentic terminal coding, rises from 78.0% to 85.8%. Terminal-bench 3.0, a harder general agent capability eval, goes from 5.4% to 14.9%. Even on that harder benchmark, the absolute number is low, but the relative improvement is nearly 3x.

For knowledge work, the GDP.pdf benchmark, which tests complex document comprehension, moves from 22.0% to 34.0%. The Harvey LAB-AA benchmark for complex legal workflows improves from 85.1% to 90.7%. These are not marginal gains. They are the kind of jumps you usually see between major model versions, not between point releases three weeks apart.

Where does 3.7 Flash land against Claude Sonnet 5 and GPT-5.6 Terra?

This is where the picture gets interesting. The model card includes comparison columns for Claude Sonnet 5, GPT-5.6 Terra, and Muse Spark 1.2, and the results are not a clean sweep for Google.

On FrontierCode 1.1 Main, 3.7 Flash leads the field at 43.6%, edging Claude Sonnet 5 at 42.7% and GPT-5.6 Terra at 41.3%. But on DeepSWE v1.1, GPT-5.6 Terra dominates at 69.6%, well ahead of 3.7 Flash at 65.3%. Claude Sonnet 5 scores 53.8% and Muse Spark 1.2 scores 54.9% on that same benchmark. The story depends on which coding benchmark you trust: if your workload looks like FrontierCode, 3.7 Flash wins. If it looks like DeepSWE, GPT-5.6 Terra is ahead.

Grouped bar chart comparing FrontierCode 1.1 Main and DeepSWE v1.1 scores across four models. Gemini 3.7 Flash: FrontierCode 43.6%, DeepSWE 65.3%. Gemini 3.6 Flash: FrontierCode 34.4%, DeepSWE 48.6%. Claude Sonnet 5: FrontierCode 42.7%, DeepSWE 53.8%. GPT-5.6 Terra: FrontierCode 41.3%, DeepSWE 69.6%.
FrontierCode 1.1 Main and DeepSWE v1.1 benchmark scores across four models. Gemini 3.7 Flash leads FrontierCode at 43.6% but GPT-5.6 Terra leads DeepSWE v1.1 at 69.6%. Source: Google DeepMind model card.

On the Artificial Analysis Intelligence Index, a composite metric, 3.7 Flash scores 56, just behind GPT-5.6 Terra and Muse Spark 1.2, which both score 57. Claude Sonnet 5 scores 55. These are close enough that the composite is not decisive.

The price gap is where 3.7 Flash separates itself. At the introductory rate of $0.75 per million input tokens and $3.75 per million output tokens, 3.7 Flash costs roughly one-third of Claude Sonnet 5, which charges $2.00 input and $10.00 output. GPT-5.6 Terra charges $2.00 input and $12.00 output. Context caching during the introductory period costs $0.075 per million tokens, compared to standard rates that will rise to $0.15 after December 31. If your agent pipeline burns output tokens through tool calls and multi-step reasoning, the cost difference compounds fast. A pipeline generating 50 million output tokens per month would cost roughly $187.50 on 3.7 Flash versus $500 on Claude Sonnet 5 and $600 on GPT-5.6 Terra at their listed rates.

What does the pricing trap mean for your agent cost model?

The introductory price is the hook. It expires on December 31, 2026, and the standard rates that take effect January 1, 2027, are exactly double: $1.50 per million input tokens and $7.50 per million output tokens. Context caching doubles too, from $0.075 to $0.15 per million tokens. Google confirmed these standard prices in its launch materials.

If you are building an agent pipeline and you size your infrastructure and pricing around the introductory rate, you are building on sand. The cost per million output tokens goes from $3.75 to $7.50 in five months. For a workload generating 100 million output tokens monthly, that is the difference between $375 and $750 per month, before you factor in input tokens and caching. Google is not hiding this: the model card and blog post both state the expiration date clearly. But anyone who has watched cloud provider free tiers and introductory credits knows how this plays out in practice. Teams build around the cheap price, ship to production, and then eat the increase because switching models mid-flight is painful.

This matters even more for agent workloads than for simple chat. Agents loop: they call tools, read results, generate plans, call more tools. Each iteration burns both input and output tokens, and the context grows with every step. A coding agent that takes 20 turns to resolve an issue might consume 500K input tokens and 50K output tokens in a single session. At the introductory rate, that session costs about $0.56. At the standard rate, it costs $1.12. If you are running thousands of sessions per day, that difference is your runway.

Should you switch your coding or agent stack to 3.7 Flash now?

The honest answer depends on what you are running today and how much token cost matters to your unit economics. Here is the breakdown:

  • If you are on Claude Sonnet 5 for coding: 3.7 Flash matches or beats it on FrontierCode and costs roughly one-third as much on output tokens. The risk is that DeepSWE v1.1 favors GPT-5.6 Terra, so if your workload is heavy on long-horizon multi-file engineering, you should benchmark both before committing. Run your own eval suite against 3.7 Flash and compare the cost per successful resolution.
  • If you are on GPT-5.6 Terra: 3.7 Flash loses on DeepSWE but wins on FrontierCode and costs about one-third on output. If your workload is more about generating production-ready code snippets and less about long-horizon repo-wide changes, the cost savings could be significant. The agentic coding benchmarks we covered previously suggest that autonomy and cost per successful task matter more than raw model intelligence for most production pipelines.
  • If you are building agent pipelines: the AutomationBench jump from 17.0% to 30.4% is the most relevant number. That benchmark tests real-world business workflow automation, which is closer to what most agent builders are doing than pure coding benchmarks. A near-doubling on that metric, combined with the price cut, makes 3.7 Flash worth a serious eval for any agent stack. Model your costs at the January 2027 rates, not the introductory ones.
  • If you are on 3.6 Flash: switch. The benchmark improvements are large enough that staying on 3.6 Flash only makes sense if you have a specific compatibility issue. The price is the same during the introductory period.

One caveat: all the benchmarks here come from Google's own model card. Google has a track record of selecting benchmarks that flatter its models. Before you migrate a production pipeline, run your own evals on your own data. The FrontierCode and AutomationBench results are impressive, but they are Google's numbers on Google's chosen benchmarks. A recent shadow evaluation of research agents showed that benchmark performance and real-world performance can diverge sharply, and the same principle applies here.

The pricing clock is the real story

The benchmarks are real and the improvements are genuine. The thing that should keep a builder up at night is whether the price they are building their business on will exist in five months. Google is offering a genuinely competitive model at a genuinely aggressive price. They are also telling you, in plain text, that the price doubles on January 1, 2027. The smart move is to treat the introductory rate as a discount and size your agent costs at $1.50 and $7.50. If the math still works, the next four months are a window to build and ship at half price. If the math only works at the introductory rate, you are building a dependency that will hurt when the bill doubles.

Sources