Skip to main content

Benchmark profile

Pencil Puzzle Bench

A multi-step verifiable reasoning benchmark that evaluates whether models can solve pencil puzzles with unique solutions.

How BenchLM shows Pencil Puzzle Bench

BenchLM mirrors the public Pencil Puzzle Bench leaderboard from July 21, 2026 snapshot. The source benchmark evaluates 51 frontier models on 300 curated puzzles spanning 20 puzzle types, with direct-ask and agentic solve rates reported separately.

Pencil Puzzle Bench is display only on BenchLM. It is a useful multi-step reasoning reference, but the public table mixes direct prompting and agentic runs and exposes variant-specific reasoning settings, so BenchLM keeps it out of weighted model rankings for now.

73 model variants300 evaluation puzzles20 puzzle types17,000 eval runsDisplay only

Best solve rate on Pencil Puzzle Bench — July 21, 2026 snapshot

BenchLM mirrors the published best solve rate view for Pencil Puzzle Bench. Claude Fable 5 leads the public snapshot at 97.6% , followed by GPT-5.5 (83.3%) and GPT-5.4 (70.2%). BenchLM does not use these results to rank models overall.

73 modelsReasoningCurrentDisplay onlyUpdated July 21, 2026 snapshot

Best solve rate table (73 models)

Score
1
Claude Fable 5Anthropic · Closed
97.6%
2
GPT-5.5OpenAI · Closed
83.3%
3
GPT-5.4OpenAI · Closed
70.2%
4
GPT-5.2OpenAI · Closed
56.0%
5
Claude Opus 4.7Anthropic · Closed
50.0%
6
Gemini 3.5 FlashGoogle · Closed
41.9%
7
Qwen3.7 MaxAlibaba · Closed
40.0%
8
GPT-5.2OpenAI · Closed
36.7%
9
36.7%
10
Claude Opus 4.6 (Adaptive)Anthropic · Closed
33.3%
11
Gemini 3.1 ProGoogle · Closed
33.3%
12
Claude Opus 4.6Anthropic · Closed
30.0%
13
GLM-5.2Z.AI · Open weight
26.7%
14
Claude Sonnet 4.6Anthropic · Closed
26.7%
15
GPT-5.2 ProOpenAI · Closed
26.7%
16
GPT-5.2OpenAI · Closed
23.3%
17
23.3%
18
23.3%
19
Kimi K2.6Moonshot AI · Open weight
20.0%
20
Kimi K2.7 CodeMoonshot AI · Open weight
16.7%
21
Qwen3.7 PlusAlibaba · Closed
16.7%
22
Gemini 3 ProGoogle · Closed
16.7%
23
Claude Sonnet 4.6Anthropic · Closed
16.7%
24
Gemini 3 ProGoogle · Closed
13.3%
25
Gemini 3 ProGoogle · Closed
10.0%
26
GPT-5.2OpenAI · Closed
10.0%
27
Qwen3.6 PlusAlibaba · Closed
10.0%
28
GPT-5.1OpenAI · Closed
7.7%
29
MiniMax M3MiniMax · Open weight
7.1%
30
Claude Opus 4.5 ThinkingAnthropic · Closed
6.7%
31
Gemini 3 FlashGoogle · Closed
6.7%
32
Gemini 3 FlashGoogle · Closed
6.7%
33
Grok 4.20xAI · Closed
6.7%
34
GPT-5 (high)OpenAI · Closed
6.0%
35
Kimi K2.5Moonshot AI · Open weight
6.0%
36
Grok 4.1 FastxAI · Closed
5.7%
37
5.3%
38
DeepSeek V4 ProDeepSeek · Open weight
4.0%
39
Grok 4.3xAI · Closed
3.3%
40
o3OpenAI · Closed
3.3%
42
MiniMax M2.5MiniMax · Closed
3.3%
43
Claude Opus 4.5Anthropic · Closed
3.3%
44
Claude Sonnet 4.5Anthropic · Closed
3.3%
46
Claude Sonnet 4.5 ThinkingAnthropic · Closed
2.3%
47
DeepSeek V3.2DeepSeek · Open weight
2.0%
48
Grok 4.3xAI · Closed
2.0%
49
Kimi K2Moonshot AI · Closed
1.3%
50
MiMo-V2-ProXiaomi · Closed
1.0%
51
o1OpenAI · Closed
0.7%
52
MiniMax M2.7MiniMax · Open weight
0.7%
53
0.7%
54
GLM-5Z.AI · Open weight
0.7%
55
Gemini 2.5 ProGoogle · Closed
0.3%
56
GPT-5.2OpenAI · Closed
0.3%
57
0.3%
58
0.3%
59
GPT-OSS 120BOpenAI · Open weight
0.3%
63
MiMo-V2-FlashXiaomi · Open weight
0.3%
64
GLM-4.7Z.AI · Open weight
0.3%
65
Grok Code Fast 1xAI · Closed
0.3%
66
Gemini 3.5 FlashGoogle · Closed
0.0%
67
0.0%
68
GPT-4.1OpenAI · Closed
0.0%
69
GPT-4oOpenAI · Closed
0.0%
70
0.0%
71
0.0%
72
0.0%
73
0.0%

The published Pencil Puzzle Bench snapshot places Claude Fable 5 first at 97.6%. The third row is 27.4 points behind. The broader top-10 range is 64.3 points, so the table still separates the published systems.

73 models have been evaluated on Pencil Puzzle Bench. The benchmark falls in the Reasoning category. This category carries a 17% weight in BenchLM.ai's overall scoring system. Pencil Puzzle Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Pencil Puzzle Bench

Year

2026

Tasks

300 evaluation puzzles

Format

Direct and agentic puzzle solve rate

Difficulty

Multi-step verifiable reasoning

BenchLM mirrors the public Pencil Puzzle Bench leaderboard as a display-only reasoning benchmark. The public site reports direct-ask and agentic solve rates across a 300-puzzle evaluation selection from the 62,231-puzzle dataset.

BenchLM freshness & provenance

Version

Pencil Puzzle Bench 2026

Refresh cadence

Static

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Pencil Puzzle Bench measure?

A multi-step verifiable reasoning benchmark that evaluates whether models can solve pencil puzzles with unique solutions.

Which model leads the published Pencil Puzzle Bench snapshot?

Claude Fable 5 currently leads the published Pencil Puzzle Bench snapshot with 97.6% best solve rate. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Pencil Puzzle Bench?

73 AI models are included in BenchLM's mirrored Pencil Puzzle Bench snapshot, based on the public leaderboard captured on July 21, 2026 snapshot.

Last updated: July 21, 2026 snapshot · mirrored from the public benchmark leaderboard

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.