Skip to main content

Benchmark profile

Artificial Analysis Coding Index (AA Coding Index)

A display-only Artificial Analysis coding index.

Data verified

Benchmark score on AA Coding Index — July 23, 2026

BenchLM mirrors the published score view for AA Coding Index. GPT-5.6 Sol leads the public snapshot at 77.4% , followed by GPT-5.6 Terra (76.7%) and Claude Fable 5 (76.5%). BenchLM does not use these results to rank models overall.

85 modelsCodingCurrentDisplay onlyUpdated July 23, 2026

Benchmark score table (85 models)

Score
1
GPT-5.6 SolOpenAI · Closed
77.4%
2
GPT-5.6 TerraOpenAI · Closed
76.7%
3
Claude Fable 5Anthropic · Closed
76.5%
4
Kimi K3Moonshot AI · Closed
76.2%
5
GPT-5.5OpenAI · Closed
74.9%
6
Claude Opus 4.8Anthropic · Closed
74.3%
7
Claude Opus 4.7 (Adaptive)Anthropic · Closed
73.6%
8
Grok 4.5xAI · Closed
72.5%
9
Claude Sonnet 5Anthropic · Closed
71.5%
10
GPT-5.6 LunaOpenAI · Closed
71.5%
11
Muse Spark 1.1Meta · Closed
71.3%
12
GPT-5.4OpenAI · Closed
71.0%
13
Gemini 3.5 FlashGoogle · Closed
70.1%
14
Gemini 3.6 FlashGoogle · Closed
69.2%
15
Gemini 3.1 ProGoogle · Closed
68.8%
16
GLM-5.2Z.AI · Open weight
68.8%
17
Qwen3.7 MaxAlibaba · Closed
66.0%
18
Kimi K2.6Moonshot AI · Open weight
61.8%
19
Kimi K2.7 CodeMoonshot AI · Open weight
60.8%
20
MiMo-V2.5-ProXiaomi · Closed
60.2%
21
DeepSeek V4 Pro (Max)DeepSeek · Open weight
59.4%
22
Hy3 PreviewTencent · Open weight
58.8%
23
Hy3Tencent · Open weight
58.8%
24
DeepSeek V4 Pro (High)DeepSeek · Open weight
58.7%
25
Muse SparkMeta · Closed
58.6%
26
MiniMax M3MiniMax · Open weight
58.6%
27
DeepSeek V4 Flash (Max)DeepSeek · Open weight
56.2%
28
GPT-5.4 miniOpenAI · Closed
56.1%
29
GPT-5.4 nanoOpenAI · Closed
56.1%
30
Qwen3.7 PlusAlibaba · Closed
55.9%
31
GLM-5.1Z.AI · Open weight
55.8%
32
Qwen3.6 PlusAlibaba · Closed
54.5%
33
Qwen3.6-27BAlibaba · Open weight
53.7%
34
MiniMax M2.7MiniMax · Open weight
52.6%
35
InklingThinking Machines Lab · Open weight
52.1%
36
DeepSeek V4 Flash (High)DeepSeek · Open weight
52.0%
37
MiMo-V2-FlashXiaomi · Open weight
49.8%
38
GPT-5.1OpenAI · Closed
49.4%
39
Gemini 3.5 Flash-LiteGoogle · Closed
49.3%
40
Nemotron 3 UltraNVIDIA · Open weight
49.3%
41
Qwen3.5 397BAlibaba · Open weight
48.2%
42
Qwen3.5 397B (Reasoning)Alibaba · Open weight
48.2%
43
Mistral Medium 3.5 128BMistral · Open weight
46.9%
44
Kimi K2.5Moonshot AI · Open weight
46.8%
45
Kimi K2.5 (Reasoning)Moonshot AI · Closed
46.8%
46
Qwen3.5-122B-A10BAlibaba · Open weight
45.7%
47
GLM-4.7Z.AI · Open weight
45.3%
48
Gemma 4 31BGoogle · Open weight
43.4%
49
Grok 4.3xAI · Closed
42.3%
50
Qwen3.6-35B-A3BAlibaba · Open weight
41.9%
51
o1OpenAI · Closed
39.7%
52
Step 3.7 FlashStepFun · Open weight
39.6%
53
Gemma 4 26B A4BGoogle · Open weight
39.3%
54
GPT-5 (high)OpenAI · Closed
37.8%
55
Nemotron 3 Super 120B A12BNVIDIA · Open weight
37.7%
56
Gemini 3.1 Flash-LiteGoogle · Closed
34.7%
57
o1-previewOpenAI · Closed
34.0%
58
Gemini 2.5 ProGoogle · Closed
33.3%
59
Mercury 2Inception · Closed
31.1%
60
GPT-OSS 120BOpenAI · Open weight
30.4%
61
Command A+Cohere · Open weight
27.9%
62
Mistral Small 4Mistral · Open weight
26.6%
63
Mistral Small 4 (Reasoning)Mistral · Open weight
26.6%
64
Ling 2.6 FlashInclusionAI · Open weight
25.3%
65
Gemini 1.5 ProGoogle · Closed
23.6%
66
DeepSeek V3DeepSeek · Open weight
23.0%
67
GPT-4 TurboOpenAI · Closed
21.5%
68
GPT-OSS 20BOpenAI · Open weight
20.7%
69
GPT-4.1 miniOpenAI · Closed
20.2%
70
Mistral Large 3Mistral · Closed
20.1%
71
Claude 3 OpusAnthropic · Closed
19.5%
72
Llama 4 MaverickMeta · Open weight
16.3%
73
GPT-5 miniOpenAI · Closed
15.6%
74
Nemotron 3 Nano 30BNVIDIA · Open weight
14.4%
75
Ministral 3 14B (Reasoning)Mistral · Open weight
14.4%
76
Ministral 3 14BMistral · Open weight
14.4%
77
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
13.8%
78
GPT-4o miniOpenAI · Closed
11.4%
79
GPT-4.1 nanoOpenAI · Closed
11.1%
80
Gemma 3 27BGoogle · Open weight
10.1%
81
Ministral 3 8B (Reasoning)Mistral · Open weight
9.7%
82
Ministral 3 8BMistral · Open weight
9.7%
83
Llama 4 ScoutMeta · Open weight
8.2%
84
Ministral 3 3B (Reasoning)Mistral · Open weight
4.8%
85
Ministral 3 3BMistral · Open weight
4.8%

The published AA Coding Index snapshot places GPT-5.6 Sol first at 77.4%. The third row is 0.9 points behind. The broader top-10 range is 5.9 points, so many of the published results sit in a relatively narrow band.

85 models have been evaluated on AA Coding Index. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. AA Coding Index is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About AA Coding Index

Year

2026

Tasks

Cross-benchmark coding index

Format

Aggregated model score

Difficulty

Display-only external reference

BenchLM mirrors this coding index for comparison, but does not use it as a weighted coding benchmark row.

BenchLM freshness & provenance

Version

AA Coding Index 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does AA Coding Index measure?

A display-only Artificial Analysis coding index.

Which model scores highest on AA Coding Index?

GPT-5.6 Sol by OpenAI currently leads with a score of 77.4% on AA Coding Index.

How many models are evaluated on AA Coding Index?

85 AI models have been evaluated on AA Coding Index on BenchLM.

Last updated: July 23, 2026 · BenchLM version AA Coding Index 2026

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.