Skip to main content

Model profile

Gemini 3.5 Flash

GoogleSupersededReleased May 19, 2026
Data verified
Superseded:Google has released newer models in this line —Gemini 3.6 Flash
Overall Score
64.75Public #33 of 200
Arena Elo
1476
Eligible category ranks
6of 8
Price (1M tokens)
$1.5 in / $9 out
API pricing
Speed
284.2tok/s
Context
1M

Evidence coverage

46 of 323 tracked benchmarks are published. 22 are verified and 24 provisional. 7 of 8 categories are measured.

Updated July 23, 2026Methodology
Published / tracked
46 / 323
Verified
22
Provisional
24
Categories with evidence
7 / 8

Evidence by category

  • Agentic15 benchmarks
    Mixed evidence
  • Coding8 benchmarks
    Mixed evidence
  • Reasoning5 benchmarks
    Mixed evidence
  • Knowledge9 benchmarks
    Mixed evidence
  • Math2 benchmarks
    Verified
  • Multilingual0 benchmarks
    Not measured
  • Multimodal5 benchmarks
    Mixed evidence
  • Inst. Following2 benchmarks
    Reported
ProprietaryReasoning
Confidence:
Very high
base

Gemini 3.5 Flash ranks #33 out of 200 models on the public leaderboard with an overall score of 64.75/100. It does not yet have enough sourced coverage for BenchLM's verified leaderboard. While not a frontier model, it offers specific advantages depending on the use case.

Gemini 3.5 Flash is a proprietary model with a 1M token context window. It uses explicit chain-of-thought reasoning, which typically improves performance on math and complex reasoning tasks at the cost of higher latency and token usage.

BenchLM links it directly to Gemini 3 Flash as the earlier related model in that lineage. This profile currently has 46 of 323 tracked benchmarks. BenchLM only exposes non-generated benchmark rows publicly, so missing categories stay blank until a sourced evaluation is available.

Its strongest category is Reasoning (#1), while its weakest is Agentic (#79). This performance profile makes it particularly strong for complex reasoning, multi-step problem solving, and analytical tasks.

Peer position

Exact provisional scores and ranks for the closest listed peers. A score can appear before a model clears the evidence threshold for a rank, so equal scores can have different rank states.

Range 64.1865.32

  1. Claude Sonnet 5
    Anthropic
    #2965.32
    Claude Sonnet 5 is #29 with a score of 65.32.
    Compare
  2. Qwen3.6 Plus
    Alibaba
    #3065.2
    Qwen3.6 Plus is #30 with a score of 65.2.
    Compare
  3. Grok 4.3
    xAI
    #3165.1
    Grok 4.3 is #31 with a score of 65.1.
    Compare
  4. Claude Sonnet 4.6
    Anthropic
    #3265.07
    Claude Sonnet 4.6 is #32 with a score of 65.07.
    Compare
  5. Gemini 3.5 FlashCurrent model
    Google
    #3364.75
    Gemini 3.5 Flash is #33 with a score of 64.75.
  6. Claude Opus 4.5
    Anthropic
    #3464.22
    Claude Opus 4.5 is #34 with a score of 64.22.
    Compare
  7. Claude Opus 4.6 (Adaptive)
    Anthropic
    #3564.18
    Claude Opus 4.6 (Adaptive) is #35 with a score of 64.18.
    Compare

Category percentile

More

Relative position among models eligible for each sourced category. A higher percentile means a stronger position within that category's ranked cohort; 100 is highest.

  1. Multimodal89%
    Eligible cohort rank #4 of 29Category score 84.0
  2. Inst. Following47%
    Eligible cohort rank #17 of 31Category score 82.4
  3. Coding85%
    Eligible cohort rank #19 of 122Category score 62.0
  4. Knowledge35%
    Eligible cohort rank #34 of 52Category score 65.5
  5. Agentic34%
    Eligible cohort rank #79 of 119Category score 44.4

Category evidence

Scores and ranks appear only where this model has published benchmark evidence. Categories without displayable source records remain not measured.

Category scores, ranks, weighting, benchmark coverage, and evidence status
CategoryScore
AgenticRank #79 of 119Percentile 34thWeight 22%15 benchmarksMixed sources44.4
CodingRank #19 of 122Percentile 85thWeight 20%8 benchmarksMixed sources62.0
ReasoningRank #1 of 1Weight 17%5 benchmarksMixed sources76.0
KnowledgeRank #34 of 52Percentile 35thWeight 12%9 benchmarksMixed sources65.5
MathRank Not rankedWeight 5%2 benchmarksVerified56.0
MultilingualWeight 7%0 benchmarksNot measuredNot measured
MultimodalRank #4 of 29Percentile 89thWeight 12%5 benchmarksMixed sources84.0
Inst. FollowingRank #17 of 31Percentile 47thWeight 5%2 benchmarksReported82.4

Chatbot Arena performance

Scroll horizontally to inspect confidence intervals and vote counts.

Chatbot Arena Elo, confidence interval, and vote count by evaluation view
ViewEloConfidence intervalVotes
Text Overall1476±6.510,094
Coding1506±11.63,021
Math1518±25.4572
Instruction Following1462±10.73,420
Creative Writing1465±16.31,502
Multi-turn1489±14.61,843
Hard Prompts1492±7.96,779
Hard Prompts (English)1484±10.93,460
Longer Query1480±9.94,393

Benchmark Details

Rows below have a displayable published verification record. Each source link and provenance note remains in the page HTML while its category is closed. Source-unverified manual rows and generated rows stay hidden.

Agentic15 benchmarks
Terminal-Bench 2.0Provider exact
76.2%Weighted 38%
Source: Google / Google DeepMind: Gemini 3.5 Flash launch screenshotsProvenance: Google reports Gemini 3.5 Flash at 76.2% on Terminal-Bench 2.1 using the Terminus-2 harness. BenchLM stores it on the existing Terminal-Bench display key.
OSWorld-VerifiedProvider exact
78.4%Weighted 34%
Source: Google / Google DeepMind: Gemini 3.5 Flash launch screenshotsProvenance: Google reports Gemini 3.5 Flash at 78.4% on OSWorld-Verified.
MCP AtlasProvider exact
83.6%Display only
Source: Google / Google DeepMind: Gemini 3.5 Flash launch screenshotsProvenance: Google reports Gemini 3.5 Flash at 83.6% on MCP Atlas.
ToolathlonProvider exact
56.5%Display only
Source: Google / Google DeepMind: Gemini 3.5 Flash launch screenshotsProvenance: Google reports Gemini 3.5 Flash at 56.5% on Toolathlon.
Finance Agent v2Benchmark exact
57.9%Display only
Source: Vals AI: Finance Agent v2Provenance: Vals Finance Agent v2 reports this exact row under google/gemini-3.5-flash; BenchLM stores it on the local financeAgentV2 key.
GDPval-AAProvider exact
1349Display only
Source: Google / Google DeepMind: Gemini 3.5 Flash launch screenshotsProvenance: Google reports Gemini 3.5 Flash at 1656 Elo on GDPval-AA.
τ²-bench resultsSecondary exact

τ²-Bench Tool-Agent-User Evaluation

95.3%Display only
Source: Artificial Analysis: Gemini 3.5 FlashProvenance: Artificial Analysis reports tau2-bench at 95.3%.
GDPval-AASecondary exact

GDPval-AA normalized

42.4%Display only
Source: Artificial Analysis: Gemini 3.5 FlashProvenance: Artificial Analysis reports normalized GDPval-AA at 57.8%.
AA Agentic IndexReported

Artificial Analysis Agentic Index

37.5%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
APEX-Agents-AAReported
47.1%Display only
Source: Artificial Analysis: apex-agents-aa leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
Gert LabsBenchmark exact

Gert Labs Composite Game Benchmark

61.85%Display only
Source: Gert Labs rankingsProvenance: Gert Labs reports this composite leaderboard score in the public rankings API. BenchLM scales the source gscore from 0-1 to 0-100 and stores it as a display-only agentic benchmark.
ResearchClawBenchBenchmark exact
18.0%Display only
Source: ResearchClawBench leaderboardProvenance: ResearchClawBench reports this model as ResearchHarness (Gemini-3.5-Flash) in the official Pass@1 leaderboard. BenchLM stores the one-decimal RADS average on the local ResearchClawBench display key and excludes it from weighted rankings.
AA AutomationBenchReported

Artificial Analysis AutomationBench

42.6%Display only
Source: Artificial Analysis: automationbench-aa leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
AA EnterpriseOps-GymReported

Artificial Analysis EnterpriseOps-Gym

50.1%Display only
Source: Artificial Analysis: enterprise-ops-gym-aa leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
terminalBenchHardReported
40.9%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Coding8 benchmarks
SciCodeSecondary exact

Scientific Code Benchmark

53.1%Weighted 16%
Source: Artificial Analysis: Gemini 3.5 FlashProvenance: Artificial Analysis reports SciCode at 53.1%.
SWE-bench ProProvider exact
55.1%Weighted 10%
Source: Google / Google DeepMind: Gemini 3.5 Flash launch screenshotsProvenance: Google reports Gemini 3.5 Flash at 55.1% on SWE-Bench Pro (Public), single attempt.
Terminal-Bench 2.0Provider exact
76.2%Display only
Source: Google / Google DeepMind: Gemini 3.5 Flash launch screenshotsProvenance: Google reports Gemini 3.5 Flash at 76.2% on Terminal-Bench 2.1 using the Terminus-2 harness. BenchLM stores it on the existing Terminal-Bench display key.
Vibe Code BenchBenchmark exact

Vibe Code Bench v1.1

48.68%Display only
Source: Vals AI: Vibe Code Bench v1.1Provenance: Vals Vibe Code Bench v1.1 reports this exact row under google/gemini-3.5-flash; BenchLM stores it on the local vibeCodeBench key.
cursorBench31Benchmark exact
49.8%Display only
Source: Cursor evals: CursorBench 3.1Provenance: Cursor reports Gemini 3.5 Flash at this exact CursorBench 3.1 score on its public evals page. BenchLM stores it on the Gemini 3.5 Flash row as a display-only coding-agent benchmark.
cursorBench32Benchmark exact
48.8%Display only
Source: Cursor evals: CursorBench 3.2Provenance: Cursor reports Gemini 3.5 Flash at this exact CursorBench 3.2 score on its public evals page. BenchLM stores it on the Gemini 3.5 Flash row as a display-only coding-agent benchmark.
AA Coding IndexReported

Artificial Analysis Coding Index

70.1%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-SciCodeReported

Artificial Analysis SciCode

53.1%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Reasoning5 benchmarks
MRCRv2Provider exact
77.3%Weighted 31%
Source: Google / Google DeepMind: Gemini 3.5 Flash launch screenshotsProvenance: Google reports Gemini 3.5 Flash at 77.3% on MRCR v2 8-needle at 128K average. BenchLM stores it on the existing MRCRv2 display key.
ARC-AGI-2Provider exact

Abstraction and Reasoning Corpus for AGI v2

72.1%Weighted 31%
Source: Google / Google DeepMind: Gemini 3.5 Flash launch screenshotsProvenance: Google reports Gemini 3.5 Flash at 72.1% on ARC-AGI-2.
MRCR 1MProvider exact
26.6%Display only
Source: Google / Google DeepMind: Gemini 3.5 Flash launch screenshotsProvenance: Google reports Gemini 3.5 Flash at 26.6% on the MRCR v2 1M pointwise long-context row.
AA-LCRSecondary exact

Artificial Analysis Long Context Reasoning

69.3%Display only
Source: Artificial Analysis: Gemini 3.5 FlashProvenance: Artificial Analysis reports AA-LCR at 69.3%.
CritPtSecondary exact

Critical Physics Tasks

13.1%Display only
Source: Artificial Analysis: Gemini 3.5 FlashProvenance: Artificial Analysis reports CritPt at 13.1%.
Knowledge9 benchmarks
HLEProvider exact

Humanity's Last Exam

40.2%Weighted 45%
Source: Google / Google DeepMind: Gemini 3.5 Flash launch screenshotsProvenance: Google reports Gemini 3.5 Flash at 40.2% on Humanity's Last Exam in the launch comparison table.
GPQASecondary exact

Graduate-Level Google-Proof Q&A

92.2%Weighted 7%
Source: Artificial Analysis: Gemini 3.5 FlashProvenance: Artificial Analysis reports Gemini 3.5 Flash at 92.2% on GPQA Diamond. BenchLM stores that exact value on the weighted GPQA lane.
Artificial Analysis Intelligence IndexSecondary exact
50.2%Display only
Source: Artificial Analysis: Gemini 3.5 FlashProvenance: Artificial Analysis reports Gemini 3.5 Flash at 55.33 on the Intelligence Index.
GPQA-DSecondary exact

GPQA Diamond

92.7%Display only
Source: Vals AI: GPQA Diamond mirrorProvenance: Vals GPQA Diamond reports this exact row under google/gemini-3.5-flash; BenchLM stores it as a display-only secondary GPQA-D reference.
AA-Omniscience AccuracySecondary exact

Artificial Analysis Omniscience Accuracy

51.9%Display only
Source: Artificial Analysis: Gemini 3.5 FlashProvenance: Artificial Analysis reports AA-Omniscience Accuracy at 51.9%.
AA-Omniscience Hallucination RateSecondary exact

Artificial Analysis Omniscience Hallucination Rate

60.7%Display only
Source: Artificial Analysis: Gemini 3.5 FlashProvenance: Artificial Analysis reports AA-Omniscience hallucination rate at 60.7%. Lower is better for this row.
AA-GPQA DiamondReported

Artificial Analysis GPQA Diamond

92.2%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-HLEReported

Artificial Analysis Humanity's Last Exam

41.0%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-Omniscience IndexReported

Artificial Analysis Omniscience Index

22.7%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Math2 benchmarks
FrontierMath v2 (Tiers 1-3)Benchmark exact

FrontierMath v2 Tiers 1-3

38.966%Weighted 30%
Source: Epoch AI FrontierMath v2 leaderboardProvenance: Epoch AI reports FrontierMath v2 Tiers 1-3 at 38.966% for gemini-3.5-flash_high. BenchLM selects the highest published thinking effort for the model and stores the v2 benchmark slice separately.
FrontierMath v2 (Tier 4)Benchmark exact

FrontierMath v2 Tier 4

14.583%Weighted 10%
Source: Epoch AI FrontierMath v2 leaderboardProvenance: Epoch AI reports FrontierMath v2 Tier 4 at 14.583% for gemini-3.5-flash_high. BenchLM selects the highest published thinking effort for the model and stores the v2 benchmark slice separately.
Multimodal5 benchmarks
MMMU-ProProvider exact

Massive Multi-discipline Multimodal Understanding Pro

83.6%Weighted 45%
Source: Google / Google DeepMind: Gemini 3.5 Flash launch screenshotsProvenance: Google reports Gemini 3.5 Flash at 83.6% on MMMU-Pro with no tools.
CharXivProvider exact

CharXiv Reasoning

84.2%Weighted 25%
Source: Google / Google DeepMind: Gemini 3.5 Flash launch screenshotsProvenance: Google reports Gemini 3.5 Flash at 84.2% on CharXiv Reasoning with no tools.
Blueprint-Bench 2Provider exact
33.6%Display only
Source: Google / Google DeepMind: Gemini 3.5 Flash launch screenshotsProvenance: Google reports Gemini 3.5 Flash at 33.6% normalized score on Blueprint-Bench 2.
AA-MMMU-ProReported

Artificial Analysis MMMU-Pro

84.3%Display only
Source: Artificial Analysis: mmmu-pro leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
Design Arena WebsiteReported

Design Arena Website Elo

1285Display only
Source: OpenRouter model benchmarksProvenance: Display-only Design Arena Website Elo synced from OpenRouter model benchmark metadata. It is excluded from BenchLM weighted scoring.
Inst. Following2 benchmarks
IFBenchSecondary exact

Instruction Following Benchmark

76.3%Weighted 65%
Source: Artificial Analysis: Gemini 3.5 FlashProvenance: Artificial Analysis reports IFBench at 76.3%.
AA-IFBenchReported

Artificial Analysis IFBench

76.3%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.

Gemini 3.5 Flash Family

Base entry

Related Earlier Model

Gemini 3 Flash

Frequently Asked Questions

How does Gemini 3.5 Flash perform overall in AI benchmarks?

Gemini 3.5 Flash currently ranks #33 out of 200 models on BenchLM's provisional leaderboard with an overall score of 64.75. It is created by Google. Its published context window is 1M.

Is Gemini 3.5 Flash good for knowledge and understanding?

Gemini 3.5 Flash ranks #34 out of 52 models in knowledge and understanding benchmarks with an average score of 65.5. There are stronger options in this category.

Is Gemini 3.5 Flash good for coding and programming?

Gemini 3.5 Flash ranks #19 out of 122 models in coding and programming benchmarks with an average score of 62. There are stronger options in this category.

Is Gemini 3.5 Flash good for mathematics?

Gemini 3.5 Flash has visible benchmark coverage in mathematics, but BenchLM does not currently assign it a global category rank there.

Is Gemini 3.5 Flash good for reasoning and logic?

Gemini 3.5 Flash ranks #1 out of 1 models in reasoning and logic benchmarks with an average score of 76. It is among the top performers in this category.

Is Gemini 3.5 Flash good for agentic tool use and computer tasks?

Gemini 3.5 Flash ranks #79 out of 119 models in agentic tool use and computer tasks benchmarks with an average score of 44.4. There are stronger options in this category.

Is Gemini 3.5 Flash good for multimodal and grounded tasks?

Gemini 3.5 Flash ranks #4 out of 29 models in multimodal and grounded tasks benchmarks with an average score of 84. It is among the top performers in this category.

Is Gemini 3.5 Flash good for instruction following?

Gemini 3.5 Flash ranks #17 out of 31 models in instruction following benchmarks with an average score of 82.4. There are stronger options in this category.

Does Gemini 3.5 Flash have full benchmark coverage on BenchLM?

Not yet. Gemini 3.5 Flash currently has 46 published benchmark scores out of the 323 benchmarks BenchLM tracks. BenchLM only exposes non-generated public benchmark rows, so missing categories stay blank until a sourced evaluation is available.

What is the context window size of Gemini 3.5 Flash?

Gemini 3.5 Flash has a published context window of 1M, which determines how much text it can process in a single interaction.

Last updated: July 23, 2026 · Runtime metrics stay blank until BenchLM has a sourced snapshot.

Choose with this week’s evidence

Join 2,000+ readers for ranking moves, new releases, pricing changes, and the evidence behind them.

Free. One email per week.