Skip to main content

Benchmark profile

Design Arena Website Elo (Design Arena Website)

A display-only Design Arena website-generation Elo score surfaced on OpenRouter model benchmark pages.

Data verified

Benchmark score on Design Arena Website — July 23, 2026

BenchLM mirrors the published score view for Design Arena Website. Kimi K3 leads the public snapshot at 1386 , followed by GLM-5.2 (1340) and Claude Fable 5 (1332). BenchLM does not use these results to rank models overall.

87 modelsMultimodal & GroundedCurrentDisplay onlyUpdated July 23, 2026

Benchmark score table (87 models)

Score
1
Kimi K3Moonshot AI · Closed
1386
2
GLM-5.2Z.AI · Open weight
1340
3
Claude Fable 5Anthropic · Closed
1332
4
Claude Opus 4.7 (Adaptive)Anthropic · Closed
1325
5
Claude Opus 4.6Anthropic · Closed
1325
6
Grok 4.5xAI · Closed
1325
7
Claude Opus 4.7Anthropic · Closed
1325
8
Claude Opus 4.6 (Adaptive)Anthropic · Closed
1325
9
Claude Sonnet 5Anthropic · Closed
1314
10
Claude Sonnet 4.6Anthropic · Closed
1314
11
Kimi K2.6Moonshot AI · Open weight
1306
12
GLM-5.1Z.AI · Open weight
1305
13
Kimi K2.7 CodeMoonshot AI · Open weight
1302
14
GLM-5-TurboZ.AI · Closed
1301
15
Muse Spark 1.1Meta · Closed
1299
16
MiMo-V2.5-ProXiaomi · Closed
1298
17
Qwen3.7 MaxAlibaba · Closed
1293
18
MiMo-V2.5Xiaomi · Closed
1291
19
MiniMax M3MiniMax · Open weight
1289
20
Qwen3.7 PlusAlibaba · Closed
1288
21
Gemini 3.5 FlashGoogle · Closed
1285
22
GPT-5.5OpenAI · Closed
1282
23
Gemini 3.1 ProGoogle · Closed
1281
24
Kimi K2.5Moonshot AI · Open weight
1279
25
Kimi K2.5 (Reasoning)Moonshot AI · Closed
1279
26
GLM-5Z.AI · Open weight
1278
27
GLM-5 (Reasoning)Z.AI · Open weight
1278
28
Claude Opus 4.5Anthropic · Closed
1277
29
Claude Opus 4.5 ThinkingAnthropic · Closed
1277
30
MiniMax M2.7MiniMax · Open weight
1275
31
Claude Opus 4.8Anthropic · Closed
1270
32
DeepSeek V4 Pro (Max)DeepSeek · Open weight
1264
33
DeepSeek V4 Pro (High)DeepSeek · Open weight
1264
34
DeepSeek V4 ProDeepSeek · Open weight
1264
35
GLM-5V-TurboZ.AI · Closed
1258
36
Grok 4.20xAI · Closed
1257
37
GLM-4.7Z.AI · Open weight
1255
38
MiniMax M2.5MiniMax · Closed
1252
39
GPT-5.4OpenAI · Closed
1250
40
Qwen3.6 PlusAlibaba · Closed
1249
41
DeepSeek V4 Flash (Max)DeepSeek · Open weight
1238
42
DeepSeek V4 Flash (High)DeepSeek · Open weight
1238
43
DeepSeek V4 FlashDeepSeek · Open weight
1238
44
Hy3Tencent · Open weight
1229
45
Gemini 3 FlashGoogle · Closed
1226
46
Grok 4.3xAI · Closed
1225
47
GPT-5.2OpenAI · Closed
1224
48
GLM-4.7-FlashZ.AI · Open weight
1224
49
Claude Sonnet 4.5Anthropic · Closed
1219
50
Claude Sonnet 4.5 ThinkingAnthropic · Closed
1219
51
GPT-5.1OpenAI · Closed
1217
52
GPT-5 (high)OpenAI · Closed
1214
53
GPT-5 (medium)OpenAI · Closed
1214
54
Step 3.7 FlashStepFun · Open weight
1211
55
Claude 4.1 OpusAnthropic · Closed
1207
56
DeepSeek V3.2DeepSeek · Open weight
1204
57
DeepSeek V3.2 (Thinking)DeepSeek · Open weight
1204
58
GLM-4.5Z.AI · Closed
1200
59
Gemini 2.5 ProGoogle · Closed
1197
60
GPT-5.3 CodexOpenAI · Closed
1193
61
GPT-5.1-CodexOpenAI · Closed
1191
62
GLM-4.5-AirZ.AI · Closed
1176
63
Claude 4 SonnetAnthropic · Closed
1175
64
Trinity-Large-PreviewArcee AI · Open weight
1165
65
Trinity-Large-ThinkingArcee AI · Open weight
1165
66
GPT-5 miniOpenAI · Closed
1154
67
Claude Haiku 4.5Anthropic · Closed
1152
68
DeepSeek V3.1 (Reasoning)DeepSeek · Open weight
1152
69
DeepSeek V3.1DeepSeek · Open weight
1152
70
Claude Haiku 4.5 ThinkingAnthropic · Closed
1152
71
DeepSeek V3DeepSeek · Open weight
1150
72
Qwen3 MaxAlibaba · Closed
1148
73
Gemini 2.5 FlashGoogle · Closed
1145
74
GPT-5 nanoOpenAI · Closed
1131
75
Nemotron 3 UltraNVIDIA · Open weight
1129
76
Mistral Medium 3Mistral · Closed
1109
77
Kimi K2Moonshot AI · Closed
1080
78
GPT-4.1OpenAI · Closed
1068
79
o3OpenAI · Closed
1067
80
GPT-4.1 miniOpenAI · Closed
1027
81
Mercury 2Inception · Closed
1016
82
GPT-4.1 nanoOpenAI · Closed
1003
83
GPT-OSS 120BOpenAI · Open weight
998
84
Llama 4 MaverickMeta · Open weight
901
85
GPT-OSS 20BOpenAI · Open weight
882
86
GPT-4oOpenAI · Closed
861
87
Llama 4 ScoutMeta · Open weight
780

The published Design Arena Website snapshot places Kimi K3 first at 1386. The third row is 54 score units behind. The broader top-10 range is 72 score units, so the table still separates the published systems.

87 models have been evaluated on Design Arena Website. The benchmark falls in the Multimodal & Grounded category. This category carries a 12% weight in BenchLM.ai's overall scoring system. Design Arena Website is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Design Arena Website

Year

2026

Tasks

Website generation comparisons

Format

Elo

Difficulty

Design and website generation

OpenRouter's Grok 4.3 benchmark page reports Website at 1294 Elo, 56.5% win rate, 166.3s average generation time, and Top 13%. BenchLM stores the Elo as the benchmark value and keeps the supporting fields in the description.

BenchLM freshness & provenance

Version

Design Arena Website 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Design Arena Website measure?

A display-only Design Arena website-generation Elo score surfaced on OpenRouter model benchmark pages.

Which model scores highest on Design Arena Website?

Kimi K3 by Moonshot AI currently leads with a score of 1386 on Design Arena Website.

How many models are evaluated on Design Arena Website?

87 AI models have been evaluated on Design Arena Website on BenchLM.

Last updated: July 23, 2026 · BenchLM version Design Arena Website 2026

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.