Skip to main content

Model profile

Claude Sonnet 5

AnthropicCurrentReleased Jun 30, 2026
Data verified
Overall Score
65.32Public #29 of 200
Arena Elo
1461
Eligible category ranks
4of 8
Price (1M tokens)
$2 in / $10 out
API pricing
Speed
Not listed
Context
1M

Evidence coverage

36 of 323 tracked benchmarks are published. 16 are verified and 20 provisional. 5 of 8 categories are measured.

Updated July 23, 2026Methodology
Published / tracked
36 / 323
Verified
16
Provisional
20
Categories with evidence
5 / 8

Evidence by category

  • Agentic13 benchmarks
    Mixed evidence
  • Coding9 benchmarks
    Mixed evidence
  • Reasoning2 benchmarks
    Reported
  • Knowledge8 benchmarks
    Mixed evidence
  • Math0 benchmarks
    Not measured
  • Multilingual0 benchmarks
    Not measured
  • Multimodal4 benchmarks
    Mixed evidence
  • Inst. Following0 benchmarks
    Not measured
ProprietaryReasoning
Confidence:
Medium
base

Claude Sonnet 5 ranks #29 out of 200 models on the public leaderboard with an overall score of 65.32/100. It does not yet have enough sourced coverage for BenchLM's verified leaderboard. This places it in the mid-tier of AI models, with strengths in specific benchmark categories.

Claude Sonnet 5 is a proprietary model with a 1M token context window. It uses explicit chain-of-thought reasoning, which typically improves performance on math and complex reasoning tasks at the cost of higher latency and token usage.

BenchLM links it directly to Claude Sonnet 4.6 as the earlier related model in that lineage. This profile currently has 36 of 323 tracked benchmarks. BenchLM only exposes non-generated benchmark rows publicly, so missing categories stay blank until a sourced evaluation is available.

Its strongest category is Knowledge (#3), while its weakest is Agentic (#15). This performance profile makes it particularly effective for knowledge-intensive tasks like research, analysis, and factual Q&A.

Peer position

Exact provisional scores and ranks for the closest listed peers. A score can appear before a model clears the evidence threshold for a rank, so equal scores can have different rank states.

Range 64.7566.27

  1. Claude Opus 4.7 (Adaptive)
    Anthropic
    #2766.27
    Claude Opus 4.7 (Adaptive) is #27 with a score of 66.27.
    Compare
  2. GLM-5
    Z.AI
    #2866.06
    GLM-5 is #28 with a score of 66.06.
    Compare
  3. Claude Sonnet 5Current model
    Anthropic
    #2965.32
    Claude Sonnet 5 is #29 with a score of 65.32.
  4. Qwen3.6 Plus
    Alibaba
    #3065.2
    Qwen3.6 Plus is #30 with a score of 65.2.
    Compare
  5. Grok 4.3
    xAI
    #3165.1
    Grok 4.3 is #31 with a score of 65.1.
    Compare
  6. Claude Sonnet 4.6
    Anthropic
    #3265.07
    Claude Sonnet 4.6 is #32 with a score of 65.07.
    Compare
  7. Gemini 3.5 Flash
    Google
    #3364.75
    Gemini 3.5 Flash is #33 with a score of 64.75.
    Compare

Category percentile

More

Relative position among models eligible for each sourced category. A higher percentile means a stronger position within that category's ranked cohort; 100 is highest.

  1. Knowledge96%
    Eligible cohort rank #3 of 52Category score 91.2
  2. Multimodal82%
    Eligible cohort rank #6 of 29Category score 78.9
  3. Coding93%
    Eligible cohort rank #9 of 122Category score 68.4
  4. Agentic88%
    Eligible cohort rank #15 of 119Category score 58.7

Category evidence

Scores and ranks appear only where this model has published benchmark evidence. Categories without displayable source records remain not measured.

Category scores, ranks, weighting, benchmark coverage, and evidence status
CategoryScore
AgenticRank #15 of 119Percentile 88thWeight 22%13 benchmarksMixed sources58.7
CodingRank #9 of 122Percentile 93rdWeight 20%9 benchmarksMixed sources68.4
ReasoningWeight 17%2 benchmarksReportedScore pending
KnowledgeRank #3 of 52Percentile 96thWeight 12%8 benchmarksMixed sources91.2
MathWeight 5%0 benchmarksNot measuredNot measured
MultilingualWeight 7%0 benchmarksNot measuredNot measured
MultimodalRank #6 of 29Percentile 82ndWeight 12%4 benchmarksMixed sources78.9
Inst. FollowingWeight 5%0 benchmarksNot measuredNot measured

Chatbot Arena performance

Scroll horizontally to inspect confidence intervals and vote counts.

Chatbot Arena Elo, confidence interval, and vote count by evaluation view
ViewEloConfidence intervalVotes
Text Overall1461±6.412,645
Coding1524±10.63,543
Math1465±23.9576
Instruction Following1459±9.64,250
Creative Writing1430±13.22,167
Multi-turn1470±13.02,220
Hard Prompts1490±7.58,369
Hard Prompts (English)1490±10.03,924
Longer Query1480±8.95,561

Benchmark Details

Rows below have a displayable published verification record. Each source link and provenance note remains in the page HTML while its category is closed. Source-unverified manual rows and generated rows stay hidden.

Agentic13 benchmarks
Terminal-Bench 2.0Provider exact
80.4%Weighted 38%
Source: Anthropic: Claude Sonnet 5 system cardProvenance: Anthropic reports Claude Sonnet 5 at 80.4% on Terminal-Bench 2.1.
OSWorld-VerifiedProvider exact
81.2%Weighted 34%
Source: Anthropic: Claude Sonnet 5 system cardProvenance: Anthropic reports Claude Sonnet 5 at 81.2% on OSWorld-Verified.
BrowseCompProvider exact
84.7%Weighted 28%
Source: Anthropic: Claude Sonnet 5 system cardProvenance: Anthropic reports Claude Sonnet 5 at 84.7% on BrowseComp in the single-agent setup.
HLE w/ toolsProvider exact

Humanity's Last Exam with tools

57.4%Display only
Source: Anthropic: Claude Sonnet 5 system cardProvenance: Anthropic reports Claude Sonnet 5 at 57.4% on Humanity’s Last Exam with tools.
GDPval-AAProvider exact
1607Display only
Source: Anthropic: Claude Sonnet 5 system cardProvenance: Anthropic reports Claude Sonnet 5 at 1609 Elo on GDPval-AA v2.
AA Agentic IndexReported

Artificial Analysis Agentic Index

46.7%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
GDPval-AAReported

GDPval-AA normalized

55.4%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA BriefcaseReported

Artificial Analysis Briefcase

1388Display only
Source: Artificial Analysis: aa-briefcase leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
AA AutomationBenchReported

Artificial Analysis AutomationBench

39.2%Display only
Source: Artificial Analysis: automationbench-aa leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
AA EnterpriseOps-GymReported

Artificial Analysis EnterpriseOps-Gym

44.7%Display only
Source: Artificial Analysis: enterprise-ops-gym-aa leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
AA Harvey LABReported

Artificial Analysis Harvey LAB-AA

90.1%Display only
Source: Artificial Analysis: harvey-lab-aa leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
AA Tau3 BankingReported

Artificial Analysis Tau3-Banking

28.2%Display only
Source: Artificial Analysis: tau3-banking leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
aaTerminalBench21Reported
80.5%Display only
Source: Artificial Analysis: terminalbench-v2-1 leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
Coding9 benchmarks
SWE-bench VerifiedProvider exact

Software Engineering Benchmark Verified

85.2%Weighted 16%
Source: Anthropic: Claude Sonnet 5 system cardProvenance: Anthropic reports Claude Sonnet 5 at 85.2% on SWE-bench Verified.
SWE-bench ProProvider exact
63.2%Weighted 10%
Source: Anthropic: Claude Sonnet 5 system cardProvenance: Anthropic reports Claude Sonnet 5 at 63.2% on SWE-bench Pro.
SWE MultilingualProvider exact
78.3%Display only
Source: Anthropic: Claude Sonnet 5 system cardProvenance: Anthropic reports Claude Sonnet 5 at 78.3% on SWE-bench Multilingual.
SWE MultimodalProvider exact

SWE-bench Multimodal

28.1%Display only
Source: Anthropic: Claude Sonnet 5 system cardProvenance: Anthropic reports Claude Sonnet 5 at 28.1% on SWE-bench Multimodal.
Terminal-Bench 2.0Provider exact
80.4%Display only
Source: Anthropic: Claude Sonnet 5 system cardProvenance: Anthropic reports Claude Sonnet 5 at 80.4% on Terminal-Bench 2.1.
FrontierCode 1.1 MainBenchmark exact
42.7%Display only
Source: Cognition: FrontierCode 1.1Provenance: Cognition reports Claude Sonnet 5 at 42.7% on FrontierCode 1.1 Main, using the best xhigh effort row from the published data JSON.
cursorBench32Benchmark exact
61.5%Display only
Source: Cursor evals: CursorBench 3.2Provenance: Cursor reports Sonnet 5 Max at this exact CursorBench 3.2 score on its public evals page. BenchLM stores it on the Claude Sonnet 5 row as a display-only coding-agent benchmark.
AA Coding IndexReported

Artificial Analysis Coding Index

71.5%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-SciCodeReported

Artificial Analysis SciCode

53.6%Display only
Source: Artificial Analysis: scicode leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
Reasoning2 benchmarks
AA-LCRReported

Artificial Analysis Long Context Reasoning

70.7%Display only
Source: Artificial Analysis: artificial-analysis-long-context-reasoning leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
CritPtReported

Critical Physics Tasks

16.9%Display only
Source: Artificial Analysis: critpt leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
Knowledge8 benchmarks
HLEProvider exact

Humanity's Last Exam

57.4%Weighted 45%
Source: Anthropic: Claude Sonnet 5 system cardProvenance: Anthropic reports Claude Sonnet 5 at 57.4% on Humanity’s Last Exam with tools. BenchLM stores this on the weighted HLE lane and separately keeps the no-tools row.
HLE w/o toolsProvider exact

Humanity's Last Exam without tools

43.2%Display only
Source: Anthropic: Claude Sonnet 5 system cardProvenance: Anthropic reports Claude Sonnet 5 at 43.2% on Humanity’s Last Exam without tools.
Artificial Analysis Intelligence IndexReported
53.4%Display only
Source: Artificial Analysis: artificial-analysis-intelligence-index leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
AA-GPQA DiamondReported

Artificial Analysis GPQA Diamond

91.1%Display only
Source: Artificial Analysis: gpqa-diamond leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
AA-HLEReported

Artificial Analysis Humanity's Last Exam

39.6%Display only
Source: Artificial Analysis: humanitys-last-exam leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
AA-Omniscience IndexReported

Artificial Analysis Omniscience Index

15.3%Display only
Source: Artificial Analysis: omniscience leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
AA-Omniscience AccuracyReported

Artificial Analysis Omniscience Accuracy

38.3%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-Omniscience Hallucination RateReported

Artificial Analysis Omniscience Hallucination Rate

37.3%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Multimodal4 benchmarks
CharXivProvider exact

CharXiv Reasoning

88.3%Weighted 25%
Source: Anthropic: Claude Sonnet 5 system cardProvenance: Anthropic reports Claude Sonnet 5 at 88.3% on CharXiv Reasoning with tools.
CharXiv w/o toolsProvider exact

CharXiv Reasoning without tools

77%Display only
Source: Anthropic: Claude Sonnet 5 system cardProvenance: Anthropic reports Claude Sonnet 5 at 77.0% on CharXiv Reasoning without tools.
AA-MMMU-ProReported

Artificial Analysis MMMU-Pro

77.3%Display only
Source: Artificial Analysis: mmmu-pro leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.
Design Arena WebsiteReported

Design Arena Website Elo

1314Display only
Source: OpenRouter model benchmarksProvenance: Display-only Design Arena Website Elo synced from OpenRouter model benchmark metadata. It is excluded from BenchLM weighted scoring.

Claude Sonnet 5 Family

Base entry

Related Earlier Model

Claude Sonnet 4.6

Frequently Asked Questions

How does Claude Sonnet 5 perform overall in AI benchmarks?

Claude Sonnet 5 currently ranks #29 out of 200 models on BenchLM's provisional leaderboard with an overall score of 65.32. It is created by Anthropic. Its published context window is 1M.

Is Claude Sonnet 5 good for knowledge and understanding?

Claude Sonnet 5 ranks #3 out of 52 models in knowledge and understanding benchmarks with an average score of 91.2. It is among the top performers in this category.

Is Claude Sonnet 5 good for coding and programming?

Claude Sonnet 5 ranks #9 out of 122 models in coding and programming benchmarks with an average score of 68.4. It is among the top performers in this category.

Is Claude Sonnet 5 good for reasoning and logic?

Claude Sonnet 5 has visible benchmark coverage in reasoning and logic, but BenchLM does not currently assign it a global category rank there.

Is Claude Sonnet 5 good for agentic tool use and computer tasks?

Claude Sonnet 5 ranks #15 out of 119 models in agentic tool use and computer tasks benchmarks with an average score of 58.7. There are stronger options in this category.

Is Claude Sonnet 5 good for multimodal and grounded tasks?

Claude Sonnet 5 ranks #6 out of 29 models in multimodal and grounded tasks benchmarks with an average score of 78.9. It is among the top performers in this category.

Does Claude Sonnet 5 have full benchmark coverage on BenchLM?

Not yet. Claude Sonnet 5 currently has 36 published benchmark scores out of the 323 benchmarks BenchLM tracks. BenchLM only exposes non-generated public benchmark rows, so missing categories stay blank until a sourced evaluation is available.

What is the context window size of Claude Sonnet 5?

Claude Sonnet 5 has a published context window of 1M, which determines how much text it can process in a single interaction.

Last updated: July 23, 2026 · Runtime metrics stay blank until BenchLM has a sourced snapshot.

Choose with this week’s evidence

Join 2,000+ readers for ranking moves, new releases, pricing changes, and the evidence behind them.

Free. One email per week.