Skip to main content

Benchmark profile

Artificial Analysis Omniscience Index (AA-Omniscience Index)

A display-only Artificial Analysis factual knowledge index.

Data verified

Benchmark score on AA-Omniscience Index — July 23, 2026

BenchLM mirrors the published score view for AA-Omniscience Index. Claude Fable 5 leads the public snapshot at 40.2% , followed by Gemini 3.1 Pro (32.9%) and Claude Opus 4.8 (27.4%). BenchLM does not use these results to rank models overall.

153 modelsKnowledgeCurrentDisplay onlyUpdated July 23, 2026

Benchmark score table (153 models)

Score
1
Claude Fable 5Anthropic · Closed
40.2%
2
Gemini 3.1 ProGoogle · Closed
32.9%
3
Claude Opus 4.8Anthropic · Closed
27.4%
4
Grok 4.5xAI · Closed
26.4%
5
Claude Opus 4.7 (Adaptive)Anthropic · Closed
26.2%
6
Gemini 3.6 FlashGoogle · Closed
23.5%
7
Gemini 3.5 FlashGoogle · Closed
22.7%
8
GPT-5.6 SolOpenAI · Closed
21.7%
9
GPT-5.5OpenAI · Closed
20.1%
10
Kimi K3Moonshot AI · Closed
18.4%
11
Grok 4.3xAI · Closed
18.3%
12
Muse Spark 1.1Meta · Closed
18.0%
13
Gemini 3 ProGoogle · Closed
15.8%
14
Claude Sonnet 5Anthropic · Closed
15.3%
15
Claude Opus 4.7Anthropic · Closed
14.2%
16
Qwen3.7 MaxAlibaba · Closed
14.1%
17
Claude Opus 4.6 (Adaptive)Anthropic · Closed
13.5%
18
Claude Opus 4.5 ThinkingAnthropic · Closed
13.3%
19
Qwen 3.6 Max (preview)Alibaba · Closed
10.2%
20
GPT-5.3 CodexOpenAI · Closed
9.9%
21
GPT-5.3-Codex-SparkOpenAI · Closed
9.9%
22
Gemini 3.5 Flash-LiteGoogle · Closed
6.9%
23
Kimi K2.6Moonshot AI · Open weight
6.4%
24
GPT-5.4OpenAI · Closed
5.7%
25
GPT-5.1OpenAI · Closed
5.6%
26
MiMo-V2-ProXiaomi · Closed
4.9%
27
Muse SparkMeta · Closed
4.1%
28
GLM-5.2Z.AI · Open weight
4.0%
29
Grok 4xAI · Closed
3.8%
30
MiMo-V2.5-ProXiaomi · Closed
3.6%
31
Claude Opus 4.6Anthropic · Closed
3.5%
32
Qwen3.6 PlusAlibaba · Closed
2.7%
33
Qwen3.7 PlusAlibaba · Closed
2.4%
34
InklingThinking Machines Lab · Open weight
2.1%
35
GLM-5Z.AI · Open weight
2.0%
36
GLM-5.1Z.AI · Open weight
1.9%
37
MiniMax M3MiniMax · Open weight
1.4%
38
MiniMax M2.7MiniMax · Open weight
0.7%
39
GPT-5.6 TerraOpenAI · Closed
-0.2%
40
Nemotron 3 UltraNVIDIA · Open weight
-0.8%
41
GPT-5.2OpenAI · Closed
-1.0%
42
GPT-5.2-CodexOpenAI · Closed
-2.5%
43
Claude Sonnet 4.6Anthropic · Closed
-2.9%
44
Gemini 3 FlashGoogle · Closed
-3.6%
45
Claude Opus 4.5Anthropic · Closed
-3.9%
46
Command A+Cohere · Open weight
-4.0%
47
GPT-5.1-Codex-MaxOpenAI · Closed
-6.0%
48
GPT-5.1-CodexOpenAI · Closed
-6.0%
49
Kimi K2.5Moonshot AI · Open weight
-8.1%
50
Kimi K2.5 (Reasoning)Moonshot AI · Closed
-8.1%
51
GPT-5 (high)OpenAI · Closed
-8.1%
52
Claude 4 SonnetAnthropic · Closed
-9.2%
53
DeepSeek V4 Pro (High)DeepSeek · Open weight
-9.7%
54
DeepSeek V4 Pro (Max)DeepSeek · Open weight
-10.0%
55
GPT-5 (medium)OpenAI · Closed
-10.1%
56
o1OpenAI · Closed
-10.5%
57
GPT-4oOpenAI · Closed
-10.7%
58
Kimi K2.7 CodeMoonshot AI · Open weight
-10.7%
59
GPT-5.6 LunaOpenAI · Closed
-11.2%
60
Gemini 2.5 ProGoogle · Closed
-14.3%
61
GLM-5-TurboZ.AI · Closed
-15.1%
62
o3OpenAI · Closed
-15.3%
63
Gemini 3.1 Flash-LiteGoogle · Closed
-15.5%
64
GPT-5 miniOpenAI · Closed
-17.2%
65
Llama 3.1 405BMeta · Open weight
-17.3%
66
MiMo-V2-OmniXiaomi · Closed
-17.4%
67
Hy3 PreviewTencent · Open weight
-18.5%
68
Hy3Tencent · Open weight
-18.5%
69
GPT-5.4 miniOpenAI · Closed
-18.7%
70
GLM-5V-TurboZ.AI · Closed
-19.0%
71
Qwen3.6-27BAlibaba · Open weight
-19.8%
72
Gemma 4 E4BGoogle · Open weight
-20.0%
73
Qwen3.6-35B-A3BAlibaba · Open weight
-21.4%
74
DeepSeek V4 Flash (High)DeepSeek · Open weight
-22.3%
75
DeepSeek V4 Flash (Max)DeepSeek · Open weight
-22.9%
76
Gemma 4 E2BGoogle · Open weight
-24.0%
77
DeepSeek-R1DeepSeek · Open weight
-27.1%
78
Kimi K2Moonshot AI · Closed
-27.5%
79
GPT-5 nanoOpenAI · Closed
-27.7%
80
DeepSeek V3.1 (Reasoning)DeepSeek · Open weight
-28.4%
81
-28.4%
82
-28.7%
83
GPT-5.4 nanoOpenAI · Closed
-29.5%
84
Qwen3.5 397BAlibaba · Open weight
-29.8%
85
Qwen3.5 397B (Reasoning)Alibaba · Open weight
-29.8%
86
Mistral Small 4Mistral · Open weight
-29.9%
87
Mistral Small 4 (Reasoning)Mistral · Open weight
-29.9%
88
Mistral Medium 3Mistral · Closed
-31.5%
89
GLM-4.6Z.AI · Open weight
-31.6%
90
LFM2.5-8B-A1BLiquidAI · Open weight
-33.3%
91
Mistral Large 2Mistral · Closed
-34.0%
92
GLM-4.7Z.AI · Open weight
-34.6%
93
Grok Code Fast 1xAI · Closed
-36.0%
94
GPT-4.1OpenAI · Closed
-36.2%
95
Mistral Medium 3.5 128BMistral · Open weight
-36.3%
96
Step 3.7 FlashStepFun · Open weight
-37.5%
97
Mistral Large 3Mistral · Closed
-39.4%
98
Qwen3.5-122B-A10BAlibaba · Open weight
-39.6%
99
MiniMax M2.5MiniMax · Closed
-39.7%
100
DeepSeek V3.1DeepSeek · Open weight
-41.1%
101
DeepSeek V3DeepSeek · Open weight
-41.3%
102
Llama 4 MaverickMeta · Open weight
-41.8%
103
Qwen3.5-27BAlibaba · Open weight
-42.0%
104
Gemini 2.5 FlashGoogle · Closed
-42.0%
105
Nemotron 3 Super 120B A12BNVIDIA · Open weight
-42.1%
106
Qwen3 MaxAlibaba · Closed
-43.1%
107
Step 3.5 FlashStepFun · Open weight
-43.7%
108
Trinity-Large-PreviewArcee AI · Open weight
-44.2%
109
Trinity-Large-ThinkingArcee AI · Open weight
-44.2%
110
Gemma 4 31BGoogle · Open weight
-45.4%
111
Nemotron Ultra 253BNVIDIA · Open weight
-45.5%
112
Qwen3.5-35B-A3BAlibaba · Open weight
-46.4%
113
DeepSeek V3.2DeepSeek · Open weight
-46.7%
114
MiniMax M1 80kMiniMax · Closed
-47.4%
115
Claude 3 HaikuAnthropic · Closed
-47.6%
116
Nova ProAmazon · Closed
-47.6%
117
Gemma 4 26B A4BGoogle · Open weight
-48.1%
118
MiMo-V2-FlashXiaomi · Open weight
-48.5%
119
GPT-OSS 120BOpenAI · Open weight
-50.0%
120
GPT-4.1 miniOpenAI · Closed
-50.1%
121
Grok 4.1 FastxAI · Closed
-50.9%
122
Nemotron 3 Nano 30BNVIDIA · Open weight
-51.6%
123
Gemma 4 12BGoogle · Open weight
-51.9%
124
Mercury 2Inception · Closed
-52.3%
125
Llama 4 ScoutMeta · Open weight
-52.4%
126
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
-56.0%
127
GPT-4.1 nanoOpenAI · Closed
-56.4%
128
Phi-4Microsoft · Open weight
-56.7%
129
K-ExaoneLG AI Research · Closed
-57.9%
130
LFM2-24B-A2BLiquidAI · Closed
-59.1%
131
GLM-4.7-FlashZ.AI · Open weight
-59.3%
132
Sarvam 105BSarvam · Open weight
-59.5%
133
Solar Pro 2Upstage · Closed
-61.7%
134
Exaone 4.0 32BLG AI Research · Open weight
-62.3%
135
GLM-4.5-AirZ.AI · Closed
-62.5%
136
GPT-OSS 20BOpenAI · Open weight
-63.9%
137
Ministral 3 3B (Reasoning)Mistral · Open weight
-64.1%
138
Ministral 3 3BMistral · Open weight
-64.1%
139
Ling 2.6 FlashInclusionAI · Open weight
-65.7%
140
Gemma 3 27BGoogle · Open weight
-65.9%
141
Ministral 3 14B (Reasoning)Mistral · Open weight
-66.8%
142
Ministral 3 14BMistral · Open weight
-66.8%
143
Ministral 3 8B (Reasoning)Mistral · Open weight
-67.8%
144
Ministral 3 8BMistral · Open weight
-67.8%
145
Sarvam 30BSarvam · Open weight
-72.0%
146
Granite-4.0-350MIBM · Open weight
-72.1%
147
Granite-4.0-H-1BIBM · Open weight
-73.6%
148
LFM2.5-1.2B-InstructLiquidAI · Closed
-73.8%
149
Granite-4.0-1BIBM · Open weight
-81.8%
150
Exaone 4.0 1.2BLG AI Research · Open weight
-82.6%
151
LFM2.5-VL-1.6B-ExtractLiquidAI · Open weight
-83.9%
152
LFM2.5-1.2B-ThinkingLiquidAI · Closed
-83.9%
153
Granite-4.0-H-350MIBM · Open weight
-87.2%

The published AA-Omniscience Index snapshot places Claude Fable 5 first at 40.2%. The third row is 12.8 points behind. The broader top-10 range is 21.8 points, so the table still separates the published systems.

153 models have been evaluated on AA-Omniscience Index. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. AA-Omniscience Index is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About AA-Omniscience Index

Year

2026

Tasks

Knowledge questions

Format

Index score

Difficulty

Broad factual knowledge

BenchLM stores the AA-Omniscience index as a display-only factuality signal alongside the accuracy and hallucination-rate rows.

BenchLM freshness & provenance

Version

AA-Omniscience Index 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does AA-Omniscience Index measure?

A display-only Artificial Analysis factual knowledge index.

Which model scores highest on AA-Omniscience Index?

Claude Fable 5 by Anthropic currently leads with a score of 40.2% on AA-Omniscience Index.

How many models are evaluated on AA-Omniscience Index?

153 AI models have been evaluated on AA-Omniscience Index on BenchLM.

Last updated: July 23, 2026 · BenchLM version AA-Omniscience Index 2026

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.