Skip to main content

Benchmark profile

ReactBench v1 (ReactBench)

A coding-agent benchmark for realistic React work, with rubrics that check production concerns such as performance, accessibility, correctness, and code quality.

How we show ReactBench

We mirror the official ReactBench v1 table from Million: 55 effort variants across 51 React tasks. The source score is a weighted rubric aggregate, and solutions that miss a blocking criterion receive zero.

ReactBench tests production React work that ordinary behavior checks can miss, including performance, accessibility, and code quality. The rows also include an agent setup and reasoning-effort choice, so we preserve the published variants and costs without adding the scores to weighted coding rankings.

55 model variants17 base models51 tasksCost per rollout preservedDisplay only

ReactBench score (Pass@1) on ReactBench — July 21, 2026 snapshot

BenchLM mirrors the published reactbench score (pass@1) view for ReactBench. GPT-5.6 Terra (max) leads the public snapshot at 53.3% , followed by GPT-5.6 Sol (max) (52.5%) and GPT-5.6 Sol (xhigh) (51.0%). BenchLM does not use these results to rank models overall.

55 modelsCodingCurrentDisplay onlyUpdated July 21, 2026 snapshot

ReactBench score (Pass@1) table (55 models)

Score
1
GPT-5.6 Terra (max)OpenAI · Closed
53.3%
2
GPT-5.6 Sol (max)OpenAI · Closed
52.5%
3
GPT-5.6 Sol (xhigh)OpenAI · Closed
51.0%
4
GPT-5.6 Terra (xhigh)OpenAI · Closed
50.2%
5
Claude Fable 5 (xhigh)Anthropic · Closed
47.1%
6
GPT-5.6 Sol (high)OpenAI · Closed
46.7%
7
GPT-5.6 Luna (max)OpenAI · Closed
43.9%
8
Claude Fable 5 (high)Anthropic · Closed
42.7%
9
GPT-5.6 Sol (medium)OpenAI · Closed
42.4%
10
GPT-5.6 Terra (high)OpenAI · Closed
41.6%
11
Claude Fable 5 (low)Anthropic · Closed
41.2%
12
Claude Fable 5 (max)Anthropic · Closed
41.2%
13
GPT-5.6 Luna (xhigh)OpenAI · Closed
41.2%
14
Grok 4.5xAI · Closed
40.4%
15
Grok 4.5xAI · Closed
40.0%
16
GPT-5.6 Luna (high)OpenAI · Closed
38.8%
17
GPT-5.6 Terra (medium)OpenAI · Closed
37.3%
18
Grok 4.5xAI · Closed
36.1%
19
GPT-5.6 Sol (low)OpenAI · Closed
35.3%
20
GPT-5.6 Terra (low)OpenAI · Closed
34.9%
21
Kimi K3Moonshot AI · Closed
32.9%
22
GLM-5.2 (xhigh)Z.AI · Open weight
32.0%
23
GLM-5.2 (high)Z.AI · Open weight
31.0%
24
GLM-5.2 (low)Z.AI · Open weight
30.2%
25
Claude Opus 4.8 (xhigh)Anthropic · Closed
29.8%
26
Claude Opus 4.8 (max)Anthropic · Closed
29.4%
27
GLM-5.2 (medium)Z.AI · Open weight
28.5%
28
GLM-5.2 (max)Z.AI · Open weight
27.8%
29
Claude Sonnet 5 (high)Anthropic · Closed
27.1%
30
Claude Opus 4.8 (high)Anthropic · Closed
25.9%
31
Claude Sonnet 5 (xhigh)Anthropic · Closed
25.9%
32
Claude Opus 4.8 (medium)Anthropic · Closed
25.5%
33
Gemini 3.5 Flash (xhigh)Google · Closed
23.9%
34
Kimi K2.7 CodeMoonshot AI · Open weight
23.5%
35
Claude Sonnet 5 (medium)Anthropic · Closed
23.1%
36
GPT-5.6 Luna (medium)OpenAI · Closed
23.1%
37
Gemini 3.1 Pro (high)Google · Closed
22.7%
38
Muse Spark 1.1 (max)Meta · Closed
22.7%
39
Claude Sonnet 5 (max)Anthropic · Closed
22.4%
40
Gemini 3.5 Flash (high)Google · Closed
22.4%
41
22.4%
42
22.4%
43
Gemini 3.1 Pro (medium)Google · Closed
22.0%
44
Muse Spark 1.1 (high)Meta · Closed
21.6%
45
Composer 2.5Cursor · Closed
21.2%
46
Gemini 3.1 Pro (xhigh)Google · Closed
21.2%
47
Claude Opus 4.8 (low)Anthropic · Closed
20.8%
48
Gemini 3.1 Pro (max)Google · Closed
19.6%
49
Muse Spark 1.1 (low)Meta · Closed
18.0%
50
Gemini 3.5 Flash (low)Google · Closed
17.6%
51
Claude Sonnet 5 (low)Anthropic · Closed
17.3%
52
17.3%
53
GPT-5.6 Luna (low)OpenAI · Closed
14.5%
54
InklingThinking Machines Lab · Open weight
10.6%
55
Gemini 3.1 Pro (low)Google · Closed
10.2%

The published ReactBench snapshot places GPT-5.6 Terra (max) first at 53.3%. The third row is 2.3 points behind. The broader top-10 range is 11.7 points, so the table still separates the published systems.

55 models have been evaluated on ReactBench. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. ReactBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About ReactBench

Year

2026

Tasks

51 production React tasks

Format

Pass@1 weighted rubric score

Difficulty

Production frontend engineering

ReactBench uses 51 tasks drawn from real React repositories. Its weighted rubric score assigns zero to solutions that miss a blocking criterion. We mirror each published reasoning-effort variant and its average rollout cost, but keep the results out of weighted coding scores because the agent setup and effort level are part of the row.

BenchLM freshness & provenance

Version

ReactBench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does ReactBench measure?

A coding-agent benchmark for realistic React work, with rubrics that check production concerns such as performance, accessibility, correctness, and code quality.

Which model leads the published ReactBench snapshot?

GPT-5.6 Terra (max) currently leads the published ReactBench snapshot with 53.3% reactbench score (pass@1). BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on ReactBench?

55 AI models are included in BenchLM's mirrored ReactBench snapshot, based on the public leaderboard captured on July 21, 2026 snapshot.

Last updated: July 21, 2026 snapshot · mirrored from the public benchmark leaderboard

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.