Skip to main content

Benchmark profile

Artificial Analysis Harvey LAB-AA (AA Harvey LAB)

An independently evaluated legal-agent benchmark from Artificial Analysis.

Data verified

Benchmark score on AA Harvey LAB — July 23, 2026

BenchLM mirrors the published score view for AA Harvey LAB. Kimi K3 leads the public snapshot at 94.6% , followed by Claude Fable 5 (93.6%) and Muse Spark 1.1 (93.1%). BenchLM does not use these results to rank models overall.

19 modelsAgenticCurrentDisplay onlyUpdated July 23, 2026

Benchmark score table (19 models)

Score
1
Kimi K3Moonshot AI · Closed
94.6%
2
Claude Fable 5Anthropic · Closed
93.6%
3
Muse Spark 1.1Meta · Closed
93.1%
4
Grok 4.5xAI · Closed
92.4%
5
Claude Opus 4.8Anthropic · Closed
91.1%
6
GLM-5.2Z.AI · Open weight
91.0%
7
Claude Sonnet 5Anthropic · Closed
90.1%
8
MiniMax M3MiniMax · Open weight
88.4%
9
GPT-5.6 LunaOpenAI · Closed
87.9%
10
GPT-5.6 SolOpenAI · Closed
87.2%
11
GPT-5.5OpenAI · Closed
86.3%
12
GPT-5.6 TerraOpenAI · Closed
85.2%
13
DeepSeek V4 Pro (Max)DeepSeek · Open weight
84.4%
14
Qwen3.7 MaxAlibaba · Closed
83.4%
15
Nemotron 3 UltraNVIDIA · Open weight
81.7%
16
DeepSeek V4 Flash (Max)DeepSeek · Open weight
81.3%
17
MiMo-V2.5-ProXiaomi · Closed
73.3%
18
Mistral Medium 3.5 128BMistral · Open weight
69.1%
19
Gemini 3.1 ProGoogle · Closed
58.9%

The published AA Harvey LAB snapshot places Kimi K3 first at 94.6%. The third row is 1.5 points behind. The broader top-10 range is 7.4 points, so many of the published results sit in a relatively narrow band.

19 models have been evaluated on AA Harvey LAB. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. AA Harvey LAB is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About AA Harvey LAB

Year

2026

Tasks

Legal agent tasks

Format

Task success rate

Difficulty

Professional legal work

Stored as a display-only agentic row.

BenchLM freshness & provenance

Version

AA Harvey LAB 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does AA Harvey LAB measure?

An independently evaluated legal-agent benchmark from Artificial Analysis.

Which model scores highest on AA Harvey LAB?

Kimi K3 by Moonshot AI currently leads with a score of 94.6% on AA Harvey LAB.

How many models are evaluated on AA Harvey LAB?

19 AI models have been evaluated on AA Harvey LAB on BenchLM.

Last updated: July 23, 2026 · BenchLM version AA Harvey LAB 2026

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.