Skip to main content

Benchmark profile

Vals MedQA (MedQA)

Evaluating language model bias in medical questions.

Data verified

How BenchLM shows MedQA

BenchLM mirrors the public Vals AI MedQA leaderboard captured from https://www.vals.ai/benchmarks/medqa and updated by Vals on April 16, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

MedQA is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

95 Vals rows7 task viewspublic datasetTasks: Overall, Unbiased, Hispanic, Black, AsianDisplay only

MedQA score on MedQA — April 16, 2026

BenchLM mirrors the published medqa score view for MedQA. O1 leads the public snapshot at 96.52% , followed by GPT-5.1 (96.38%) and Gemini 3.1 Pro Preview (96.37%). BenchLM does not use these results to rank models overall.

95 modelsExternal benchmark mirrorsCurrentDisplay onlyUpdated April 16, 2026

MedQA score table (95 models)

Score
1
O1OpenAI
96.52%
2
GPT-5.1OpenAI
96.38%
4
GPT-5OpenAI
96.32%
5
GPT-5.4OpenAI
96.09%
6
O3OpenAI
96.06%
7
96.06%
9
O4 MiniOpenAI
96.02%
12
95.41%
13
95.21%
14
O3 MiniOpenAI
94.83%
16
94.55%
17
94.37%
18
94.27%
19
GPT-5.2OpenAI
94.13%
20
93.92%
21
GLM 4.7Zhipu AI
93.74%
23
93.26%
24
93.16%
26
93.01%
27
Claude Opus 4Anthropic
92.87%
29
Kimi K2 ThinkingMoonshot AI
92.59%
30
92.53%
31
92.53%
32
Grok 4 0709SpaceXAI
92.49%
33
Grok 2 1212SpaceXAI
92.32%
34
GLM 4.6Zhipu AI
92.22%
35
92.08%
36
92.07%
37
92.06%
39
GPT Oss 120bFireworks AI
91.36%
40
GPT-4.1OpenAI
91.18%
42
91.16%
44
DeepSeek R1Fireworks AI
90.80%
45
Qwen3 235b A22bFireworks AI
90.62%
46
90.35%
47
O1 MiniOpenAI
90.22%
50
GLM 4.5Zhipu AI
89.97%
51
89.47%
52
DeepSeek V3p2Fireworks AI
89.45%
56
GPT-4oOpenAI
88.16%
57
87.38%
58
Qwen3 MaxAlibaba
87.37%
61
84.63%
62
83.97%
63
Grok 3SpaceXAI
83.85%
64
83.19%
65
GPT Oss 20bFireworks AI
82.88%
66
82.36%
67
82.23%
68
DeepSeek V3 0324Fireworks AI
82.00%
69
81.99%
70
81.47%
71
DeepSeek V3Fireworks AI
80.90%
72
80.55%
75
78.23%
77
76.53%
78
76.22%
81
72.44%
82
69.10%
83
68.22%
84
68.11%
87
58.47%
88
56.98%
89
55.18%
90
53.22%
91
52.52%
93
50.70%
94
43.30%
95
2.65%

The published MedQA snapshot places O1 first at 96.52%. The third row is 0.15 points behind. The broader top-10 range is 0.64 points, so many of the published results sit in a relatively narrow band.

95 models have been evaluated on MedQA. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. MedQA is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About MedQA

Year

2026

Tasks

Medical question answering

Format

Accuracy score

Difficulty

Medical knowledge and bias evaluation

BenchLM mirrors the public Vals AI MedQA leaderboard as display-only external evidence. The captured snapshot preserves overall scores, task-level scores where Vals publishes them, uncertainty, latency, and cost-per-test metadata. It is excluded from BenchLM weighted rankings.

BenchLM freshness & provenance

Version

MedQA 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does MedQA measure?

Evaluating language model bias in medical questions.

Which model leads the published MedQA snapshot?

O1 currently leads the published MedQA snapshot with 96.52% medqa score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on MedQA?

95 AI models are included in BenchLM's mirrored MedQA snapshot, based on the public leaderboard captured on April 16, 2026.

Last updated: April 16, 2026 · mirrored from the public benchmark leaderboard

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.