Google Launches Gemini 3.7 Flash With Sharp Coding Benchmark Gains, Same-Day GitHub Copilot Rollout
Google's Gemini 3.7 Flash posts sharp coding and web-dev benchmark gains at half its predecessor's launch price, and lands in GitHub Copilot immediately.
Overview
Google released Gemini 3.7 Flash on August 13, describing it as “our most intelligent workhorse model yet for coding and agents”, and the model was available in GitHub Copilot the same day. The release came just three weeks after Gemini 3.6 Flash, continuing Google’s rapid Flash-tier release cadence, and Google Senior Director of Product Management Tulsee Doshi presented it as targeted squarely at software engineering, web development, and knowledge work.
What We Know
Google reported that “3.7 Flash shows strong gains over 3.6 Flash in coding tasks like debugging and issue resolution”, citing several benchmarks:
- On FrontierCode 1.1 Main, the score rose from 34.4% to 43.6%, and on DeepSWE v1.1, performance jumped from 49.0% to 65.3%.
- In web development, the model’s Elo score on Arena.ai’s WebDev Arena rose from 1538 to 1588, with Google saying the model generates “more functional layouts and feature-complete apps in fewer prompts”.
- The model “significantly outperforms 3.6 Flash on the GDP.pdf benchmark (34.0% vs 22.0%)”, a document-processing test, and it climbed from 17.0% to 30.4% on AutomationBench, a business-workflow benchmark.
On pricing, Google set an introductory rate of “$0.75/1M input tokens and $3.75/1M output tokens”, which 9to5Google noted is half of what 3.6 Flash launched at. That introductory pricing runs through December 31, 2026, after which standard pricing of $1.50/1M input tokens and $7.50/1M output tokens takes effect starting January 1, 2027, according to Google.
Google made the model available to developers through Google AI Studio, Google Antigravity, and Android Studio, to enterprises through the Gemini Enterprise Agent Platform, and to individual subscribers through Gemini Spark for AI Pro and Ultra plans.
GitHub rolled the model into Copilot the same day, writing that “from early testing, this model has made improvements in web and app development and agentic coding workflows over its previous version”, and that it “delivers improvements in code quality, final-output presentation, codebase research, and verification during complex coding tasks”. Inside Copilot, the model reached Pro, Pro+, Max, Business, and Enterprise plan users through a model picker in Visual Studio Code, Visual Studio, Copilot CLI, the GitHub Copilot cloud agent, the GitHub Copilot app, JetBrains, Xcode, and Eclipse, with GitHub describing the rollout as gradual and billing it under “provider list pricing under usage-based billing”. GitHub also said Copilot Enterprise and Business administrators must enable a “Gemini 3.7 Flash Preview” policy before organization members can use the model.
What We Don’t Know
Neither Google nor GitHub disclosed the underlying test sets or methodology behind FrontierCode 1.1 Main, DeepSWE v1.1, GDP.pdf, or AutomationBench beyond the headline percentage comparisons, so it isn’t possible to independently verify how representative these benchmarks are of real-world coding workloads. GitHub’s changelog entry also did not specify exactly when the gradual rollout across all listed IDEs and clients will be complete.
Analysis
Gemini 3.7 Flash’s same-day arrival in GitHub Copilot places it alongside a string of other models GitHub has added to its Copilot model picker this year, including Grok 4.6, MAI-Code-1.1-Flash, and Kimi K3, underscoring how quickly Copilot’s model lineup is turning over as providers compete on both coding benchmark scores and per-token pricing for agentic and terminal-based development work.