Skip to main content

Benchmark profile

APEX-Agents-AA

Artificial Analysis' implementation of the APEX-Agents benchmark for long-horizon professional-services agent tasks.

Data verified

Benchmark score on APEX-Agents-AA — July 23, 2026

BenchLM mirrors the published score view for APEX-Agents-AA. Gemini 3.5 Flash leads the public snapshot at 47.1% , followed by Kimi K3 (41.3%) and GPT-5.6 Terra (38.9%). BenchLM does not use these results to rank models overall.

27 modelsAgenticCurrentDisplay onlyUpdated July 23, 2026

Benchmark score table (27 models)

Score
1
Gemini 3.5 FlashGoogle · Closed
47.1%
2
Kimi K3Moonshot AI · Closed
41.3%
3
GPT-5.6 TerraOpenAI · Closed
38.9%
4
GPT-5.5OpenAI · Closed
37.7%
5
GPT-5.6 LunaOpenAI · Closed
35.8%
6
GLM-5.2Z.AI · Open weight
33.7%
7
GPT-5.4OpenAI · Closed
33.3%
8
Claude Opus 4.6 (Adaptive)Anthropic · Closed
33.0%
9
Gemini 3.1 ProGoogle · Closed
32.0%
10
Kimi K2.6Moonshot AI · Open weight
28.5%
11
GPT-5.4 miniOpenAI · Closed
28.2%
12
GPT-5.4 nanoOpenAI · Closed
24.9%
13
DeepSeek V4 Pro (Max)DeepSeek · Open weight
24.3%
14
Qwen3.7 PlusAlibaba · Closed
22.4%
15
Grok 4.3xAI · Closed
17.0%
16
Qwen3.5 397BAlibaba · Open weight
15.3%
17
Qwen3.5 397B (Reasoning)Alibaba · Open weight
15.3%
18
Step 3.7 FlashStepFun · Open weight
14.8%
19
GLM-5Z.AI · Open weight
14.5%
20
Gemini 3.1 Flash-LiteGoogle · Closed
12.2%
21
Kimi K2.5Moonshot AI · Open weight
11.5%
22
Kimi K2.5 (Reasoning)Moonshot AI · Closed
11.5%
23
MiniMax M2.7MiniMax · Open weight
10.6%
24
GPT-OSS 120BOpenAI · Open weight
3.1%
25
MiMo-V2.5-ProXiaomi · Closed
2.4%
26
Nemotron 3 Super 120B A12BNVIDIA · Open weight
1.8%
27
GPT-OSS 20BOpenAI · Open weight
0.7%

The published APEX-Agents-AA snapshot places Gemini 3.5 Flash first at 47.1%. The third row is 8.2 points behind. The broader top-10 range is 18.6 points, so the table still separates the published systems.

27 models have been evaluated on APEX-Agents-AA. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. APEX-Agents-AA is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About APEX-Agents-AA

Year

2026

Tasks

452 professional-services agent tasks

Format

Pass@1

Difficulty

Long-horizon workplace agent tasks

BenchLM stores APEX-Agents-AA as a display-only agentic row. Artificial Analysis reports pass@1 over 452 public APEX-Agents tasks spanning investment banking, management consulting, and corporate law.

BenchLM freshness & provenance

Version

APEX-Agents-AA 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does APEX-Agents-AA measure?

Artificial Analysis' implementation of the APEX-Agents benchmark for long-horizon professional-services agent tasks.

Which model scores highest on APEX-Agents-AA?

Gemini 3.5 Flash by Google currently leads with a score of 47.1% on APEX-Agents-AA.

How many models are evaluated on APEX-Agents-AA?

27 AI models have been evaluated on APEX-Agents-AA on BenchLM.

Last updated: July 23, 2026 · BenchLM version APEX-Agents-AA 2026

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.