Skip to content
FoxyPulse

VS BENCHMARK

Gemini 3.8 Flash Early Benchmark Audit: What Google Published vs Artificial Analysis (Sep 2026)

Audit dated 4 September 2026: Gemini 3.8 Flash as Google published it on 2 Sep 2026 versus Artificial Analysis Intelligence Index, speed, and task-cost. Vendor evals labeled. No invented SWE-bench or 33x cost math.

Updated
Reading time
7 min
Research desk
FoxyPulse editorial

This is an audit of published numbers, not a lab bake-off. Google announced Gemini 3.8 Flash and Gemini 3.8 Flash Cyber on 2 September 2026. The live URL slug on this page stays as-is; the model to audit is Gemini 3.8 Flash, not a generic “3.8” checkpoint. Sources checked 4 September 2026: Google’s launch blog and the Artificial Analysis (AA) Gemini 3.8 Flash release card. Vendor evals are labeled as vendor evals. Independent index scores are labeled as AA.

A same-week press table that pairs Gemini against vendor-day-one GPT-6 Astra / Claude Fable 5.1 scores is not reproduced here. A new Fable/Astra comparison page is skipped until independent primary scores exist.

What Google published (2 September 2026)

Google describes Gemini 3.8 Flash as its “most intelligent workhorse model,” with two variants on the same foundational intelligence:

  • Gemini 3.8 Flash — software engineering, agentic tasks, and multi-step reasoning in specialized domains, at the same introductory API price as Gemini 3.7 Flash.
  • Gemini 3.8 Flash Cyber — vulnerability detection and automated patching, available to trusted defenders through Google’s Fairwind Program (prioritized access for trusted government authorities, critical infrastructure operators, and software maintainers).

Introductory API price (Google): $0.75 per million input tokens and $3.75 per million output tokens, matching 3.7 Flash. Google’s footnote says that introductory price expires on 31 December 2026; from 1 January 2027 the listed rates are $1.50 / 1M input and $7.50 / 1M output.

Availability (Google): Gemini API, Google AI Studio, Antigravity, and Gemini Enterprise for developers and organizations; Google AI Pro and Ultra subscribers in the Gemini app, AI Mode in Search, and Gemini in Sheets. Gemini 3.7 Flash remains fully supported for efficiency-first workloads.

Google-reported evals (vendor, not AA)

These figures come from Google’s blog. They are not Artificial Analysis reproductions and they are not SWE-bench Verified.

  • DeepSWE v1.1 (long-horizon software engineering): Google says 3.8 Flash “outperforms most larger frontier models” at a fraction of the cost. The blog text does not publish a DeepSWE percentage. This page does not invent one.
  • HLE-Verified: 54.9%, per Google, as a multi-step reasoning check across STEM, humanities, and professional fields.
  • Vals Finance Agent V2 and Harvey’s Legal Agent Benchmark: Google says 3.8 Flash outperforms 3.7 Flash and other frontier models. No additional scores appear in that blog paragraph.
  • Effort / tokens: Google says 3.8 Flash “works harder” on complex tasks — extra reasoning steps, iterative tool calls, and at times more tokens at higher effort. Lower effort is presented as the efficiency control; 3.7 Flash remains the efficiency-first option.

Gemini 3.8 Flash Cyber — Google-reported, Fairwind-only

Flash Cyber is not the consumer default. Google limits it to the Fairwind Program.

  • CWE-Bench (Collinear, pass@1): Google reports 47.2% for 3.8 Flash Cyber versus 47.8% for a leading frontier model, at significantly lower cost. That is Google’s comparison, not an AA table.
  • Chrome Security (Google-reported of its own team): Google says the Chrome Security team found 3.8 Flash Cyber produced 2.6 times more correct patches than the best commercial models that team compared. Label: Google reporting its own internal team result.
  • CyberGym: Google describes frontier-level autonomous vulnerability discovery versus 3.5 Flash Cyber and larger frontier models. No extra CyberGym percentage is quoted here because the retrieved blog paragraph did not state one.

What Artificial Analysis published (retrieved 4 September 2026)

Artificial Analysis scores the Gemini 3.8 Flash release as three effort variants. AA Intelligence Index v4.1.1 is a weighted composite of nine evaluations: GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR. That index is not SWE-bench Verified. This page does not print a SWE-bench Verified score for Gemini 3.8 Flash.

AA variant Intelligence Index Output speed Cost per Intelligence Index task
Gemini 3.8 Flash (high) 59 327 t/s $0.58
Gemini 3.8 Flash (medium) 57 312 t/s $0.41
Gemini 3.8 Flash (low) 52 313 t/s $0.24
  • Context window (AA table): 1M.
  • Lowest time-to-first-token (AA): Gemini 3.8 Flash (low) at 0.70s. This page does not cite 142 ms.
  • AA cost column is cost per Intelligence Index task, not Google’s $ / 1M token API list price. Do not divide one by the other to invent a “33x” agent-loop saving.

How to read Google vs AA

Google is arguing that a Flash-class model can approach higher-cost frontier work on selected vendor evals (DeepSWE qualitative claim, HLE-Verified 54.9%, finance/legal agent benches without extra published scores). AA is measuring a different object: an intelligence index plus speed and task-cost at three effort levels. Those two records can both be true and still not be a head-to-head.

Use Google’s API list price if you are budgeting tokens. Use AA’s task-cost and t/s if you are comparing AA’s own index. Do not mix Google HLE-Verified with AA Intelligence Index into a single fake ranking, and do not backfill SWE-bench Verified from another model family.

What this audit does not claim

  • No 142 TPS, no 74.2% SWE-bench Verified, no 142 ms TTFT, and no 33x cost-advantage math — those figures were not in the Google blog or the AA release card retrieved for this snapshot.
  • No Claude Opus 5 / GPT-5.6 Sol / Grok 4.6 comparison table. Those rows had no primary cite on the previous version of this page.
  • No OpenAI GPT-6 Astra or Claude Fable 5.1 reprint from launch-day press.

Drawbacks, Limitations & Risks (Cons)

  • Token pricing and rate limit volatility: High-concurrency enterprise API workloads can encounter sudden RPM/TPM rate limits or regional inference tier throttling during peak developer hours.
  • Context window degradation and retrieval drift: While extended context windows allow processing 128k+ to 1M tokens, needle-in-a-haystack recall fidelity declines slightly past the 75% context threshold.
  • Self-hosting VRAM requirements: Running full unquantized FP16 weights locally requires substantial multi-GPU hardware (dual RTX 4090s or Apple M-series Max chips with 64GB+ unified memory).

Recommended API Routing & Fallback Infrastructure

Production engineering teams deploying frontier reasoning models maintain automated fallback tiers to protect uptime against upstream provider rate limits and regional outages. Configuring dynamic multi-model routing via Cloudways provides immediate failover across providers without SDK rewrites, while serverless inference platforms like Cloudways and low-latency providers like Groq deliver ultra-fast time-to-first-token (TTFT) when deploying open-weights fallbacks.

FAQ

Is Gemini 3.8 the same as Gemini 3.8 Flash?

Google’s 2 September 2026 post introduces Gemini 3.8 Flash and Gemini 3.8 Flash Cyber. This audit treats Flash as the public workhorse and Cyber as the Fairwind-gated variant. The URL slug is unchanged.

How does 3.8 Flash compare to 3.7 Flash on Google’s own page?

Google says 3.8 Flash improves on 3.7 Flash for software engineering, agentic tasks, and specialized multi-step reasoning, at the same introductory price, while 3.7 Flash stays supported for efficiency-first work. Google does not publish a SWE-bench Verified percentage for that comparison in the retrieved blog text.

What is the independent speed number?

Artificial Analysis, retrieved 4 September 2026: output speed 327 t/s (high), 312 t/s (medium), 313 t/s (low); lowest TTFT 0.70s on the low-effort variant.

Primary sources (retrieved 4 September 2026): Google blog — Introducing Gemini 3.8 Flash and 3.8 Flash Cyber (2 September 2026); Artificial Analysis — Gemini 3.8 Flash release card. Fable/Astra launch coverage is noted only as a reason a comparison page was skipped, not as a score source.