Skip to main content

Benchmark profile

LisanBench

A word-chain reasoning benchmark that tests planning, recall, constraint following, and vocabulary depth by asking models to extend non-repeating edit-distance-1 chains.

How BenchLM shows LisanBench

BenchLM mirrors the current public LisanBench difficulty-weighted leaderboard using the official dataset published at lisanbench.com for the July 21, 2026 snapshot. The public benchmark tests 150 model variants across 50 starting words, with 3 trials per starting word.

LisanBench is a strong reasoning reference, but BenchLM currently keeps it display only rather than weighted. The public leaderboard is highly variant-specific, strongly English-vocabulary-dependent, and not yet aligned cleanly enough with BenchLM canonical model rows to use as a ranking input.

150 model variants50 starting words3 trials per wordDifficulty-weighted scoresDisplay only

Difficulty-weighted score on LisanBench — July 21, 2026 snapshot

BenchLM mirrors the published difficulty-weighted score view for LisanBench. Opus 4.7 (xhigh) leads the public snapshot at 5122.60 , followed by Claude Fable 5 (medium) (4561.82) and Opus 4.6 (16k) (3526.49). BenchLM does not use these results to rank models overall.

150 modelsReasoningCurrentDisplay onlyUpdated July 21, 2026 snapshot

Difficulty-weighted score table (150 models)

Score
1
Opus 4.7 (xhigh)Anthropic · Closed
5122.60
2
Claude Fable 5 (medium)Anthropic · Closed
4561.82
3
3526.49
4
GPT 5.5 (medium)OpenAI · Closed
3315.52
5
2944.27
6
GPT 5.4 (medium)OpenAI · Closed
2738.16
7
Opus 4.8 (high)Anthropic · Closed
2693.67
8
2204.43
9
1929.11
10
Grok 4 (medium)xAI · Closed
1778.40
11
Sonnet 5 (high)Anthropic · Closed
1736.56
12
O3 (medium)OpenAI · Closed
1523.00
14
1464.63
15
GPT 5.2 (medium)OpenAI · Closed
1458.80
16
GPT 5 (medium)OpenAI · Closed
1457.28
17
GPT 5.6 Sol (medium)OpenAI · Closed
1403.28
18
GPT 5.6 Terra (medium)OpenAI · Closed
1196.90
19
1130.63
20
Gemini 3.5 Flash (high)Google · Closed
1128.26
21
1090.82
22
Deepseek V4 Flash (high)DeepSeek · Open weight
1063.47
23
Deepseek V4 Pro (high)DeepSeek · Open weight
1059.51
24
Deepseek V3.2 (thinking)DeepSeek · Open weight
925.31
25
872.67
26
Step 3.5 Flash (thinking)StepFun · Open weight
811.21
28
GPT 5 Mini (medium)OpenAI · Closed
758.84
29
GPT 5.6 Luna (medium)OpenAI · Closed
648.18
30
Kimi K2.5 (thinking)Moonshot AI · Closed
641.96
31
Kimi K2 (thinking)Moonshot AI · Closed
633.05
32
GPT 5 Nano (medium)OpenAI · Closed
626.86
33
604.43
34
602.95
35
GLM 5.2 (high)Z.AI · Open weight
591.99
36
591.88
37
GPT 5.4 Mini (medium)OpenAI · Closed
591.48
38
GPT 5.4 Nano (medium)OpenAI · Closed
543.21
39
O3 Mini (medium)OpenAI · Closed
518.37
41
GPT-OSS-120B (medium)OpenAI · Open weight
448.33
42
Qwen3.5 397B A17B (thinking)Alibaba · Open weight
387.74
43
O4 Mini (medium)OpenAI · Closed
352.68
44
GLM 5 (thinking)Z.AI · Open weight
336.34
45
332.46
46
GPT 5.5OpenAI
305.30
47
Opus 4.8Anthropic · Closed
270.10
49
Opus 4Anthropic
262.56
51
Minimax M2.5 (thinking)MiniMax · Closed
228.38
52
Qwen3 235B A22B 2507 (thinking)Alibaba · Open weight
226.22
53
Opus 4.7Anthropic · Closed
217.52
54
Opus 4.1Anthropic · Closed
215.66
55
Sonnet 4.6Anthropic · Closed
208.64
56
Sonnet 5Anthropic
208.26
57
Gemini 2.5 Pro (16k)Google · Closed
197.43
58
194.03
59
Grok 3 (thinking)xAI · Closed
188.12
60
Sonnet 3.7Anthropic
162.44
61
GPT-OSS-20B (medium)OpenAI · Open weight
156.99
62
156.14
64
Sonnet 4Anthropic · Closed
150.17
65
Sonnet 3.6Anthropic · Closed
149.65
66
Sonnet 3.5Anthropic · Closed
129.73
67
Gemini Pro 1.5Google · Closed
119.72
68
Deepseek V3.2DeepSeek · Open weight
119.04
69
Deepseek V4 ProDeepSeek · Open weight
117.14
71
Deepseek R1 0528 (thinking)DeepSeek · Open weight
111.60
72
Qwen3.5 122B A10B (thinking)Alibaba · Open weight
109.62
73
GPT 5.4OpenAI · Closed
109.51
74
GLM 4.5 (thinking)Z.AI · Closed
108.32
75
Qwen3.5 35B A3B (thinking)Alibaba · Open weight
107.61
76
104.86
77
Sonnet 4.5Anthropic · Closed
103.58
78
Deepseek V3DeepSeek · Open weight
103.39
79
103.10
80
96.63
81
GPT 4oOpenAI · Closed
94.16
82
Opus 4.5Anthropic · Closed
93.49
83
Opus 4.6Anthropic · Closed
91.61
84
GPT 4 TurboOpenAI · Closed
91.48
85
Kimi K2Moonshot AI · Closed
85.92
86
77.21
87
Opus 3Anthropic · Closed
75.77
88
Gemini 2.5 FlashGoogle · Closed
72.17
89
Minimax M1 (thinking)MiniMax · Closed
66.71
90
64.50
91
62.04
92
Deepseek V4 FlashDeepSeek · Open weight
56.78
93
Horizon BetaOpenRouter
55.69
94
55.20
95
54.86
96
Nova Pro V1Amazon · Closed
54.38
97
GLM 4.7 (thinking)Z.AI · Open weight
54.27
98
Polaris AlphaOpenRouter
53.34
99
Haiku 4.5Anthropic · Closed
52.64
100
50.96
101
50.95
102
48.88
103
Grok 4.1 FastxAI · Closed
47.29
104
GLM 4.6 (thinking)Z.AI · Open weight
44.20
105
43.79
106
Mistral Medium 3Mistral · Closed
43.31
107
42.74
108
Llama 4 MaverickMeta · Open weight
42.33
109
GPT 4.1OpenAI · Closed
42.02
110
40.99
111
40.07
112
Devstral MediumMistral AI
40.03
114
Haiku 3.5Anthropic
38.22
115
38.21
116
38.06
117
38.06
118
Haiku 3Anthropic · Closed
35.46
119
34.64
120
GPT 4.1 MiniOpenAI · Closed
32.82
122
31.92
123
Llama 4 ScoutMeta · Open weight
30.85
124
Mimo V2 Flash (thinking)Xiaomi · Open weight
28.75
125
25.24
126
24.88
127
24.38
128
Qwen3 32BAlibaba
23.86
129
21.56
130
GPT 4o MiniOpenAI · Closed
21.21
131
18.98
132
17.77
133
Qwen3 14BAlibaba
16.68
134
Qwen3 8BAlibaba
15.50
135
GPT 4.1 NanoOpenAI · Closed
14.95
136
Devstral SmallMistral AI
13.22
137
Codestral 2508Mistral AI
13.15
138
12.29
139
11.78
140
Mistral NemoMistral AI
11.47
141
11.16
142
9.41
143
9.24
144
Qwen3 4BAlibaba
7.90
145
6.64
146
Qwen3 1.7BAlibaba
6.27
147
3.92
148
2.85
149
0.63
150
Qwen3 0.6BAlibaba
0.06

The published LisanBench snapshot places Opus 4.7 (xhigh) first at 5122.60. The third row is 1596.11 score units behind. The broader top-10 range is 3344.20 score units, so the table still separates the published systems.

150 models have been evaluated on LisanBench. The benchmark falls in the Reasoning category. This category carries a 17% weight in BenchLM.ai's overall scoring system. LisanBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About LisanBench

Year

2026

Tasks

50 starting words × 3 trials

Format

Difficulty-weighted word-chain reasoning

Difficulty

Open-ended lexical planning

BenchLM mirrors the public difficulty-weighted LisanBench leaderboard as a display-only reasoning benchmark. The public benchmark currently evaluates 128 model variants across 50 starting words with 3 trials per word.

BenchLM freshness & provenance

Version

LisanBench 2026

Refresh cadence

Static

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does LisanBench measure?

A word-chain reasoning benchmark that tests planning, recall, constraint following, and vocabulary depth by asking models to extend non-repeating edit-distance-1 chains.

Which model leads the published LisanBench snapshot?

Opus 4.7 (xhigh) currently leads the published LisanBench snapshot with 5122.60 difficulty-weighted score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on LisanBench?

150 AI models are included in BenchLM's mirrored LisanBench snapshot, based on the public leaderboard captured on July 21, 2026 snapshot.

Last updated: July 21, 2026 snapshot · mirrored from the public benchmark leaderboard

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.