Skip to main content

Benchmark profile

Artificial Analysis Tau3-Banking (AA Tau3 Banking)

An independently evaluated Tau3 banking benchmark from Artificial Analysis.

Data verified

Benchmark score on AA Tau3 Banking — July 23, 2026

BenchLM mirrors the published score view for AA Tau3 Banking. Kimi K3 leads the public snapshot at 33.4% , followed by GPT-5.6 Sol (33.0%) and Grok 4.5 (32.6%). BenchLM does not use these results to rank models overall.

19 modelsAgenticCurrentDisplay onlyUpdated July 23, 2026

Benchmark score table (19 models)

Score
1
Kimi K3Moonshot AI · Closed
33.4%
2
GPT-5.6 SolOpenAI · Closed
33.0%
3
Grok 4.5xAI · Closed
32.6%
4
GPT-5.6 TerraOpenAI · Closed
31.8%
5
GPT-5.5OpenAI · Closed
31.3%
6
Claude Sonnet 5Anthropic · Closed
28.2%
7
Claude Opus 4.8Anthropic · Closed
27.6%
8
GPT-5.6 LunaOpenAI · Closed
27.2%
9
Claude Fable 5Anthropic · Closed
26.8%
10
GLM-5.2Z.AI · Open weight
26.8%
11
DeepSeek V4 Pro (Max)DeepSeek · Open weight
25.8%
12
Muse Spark 1.1Meta · Closed
25.2%
13
Gemini 3.6 FlashGoogle · Closed
24.5%
14
InklingThinking Machines Lab · Open weight
23.7%
15
DeepSeek V4 Flash (Max)DeepSeek · Open weight
22.9%
16
Gemini 3.1 ProGoogle · Closed
16.5%
17
Gemini 3.5 Flash-LiteGoogle · Closed
16.5%
18
Gemma 4 31BGoogle · Open weight
15.1%
19
Mistral Medium 3.5 128BMistral · Open weight
14.4%

The published AA Tau3 Banking snapshot places Kimi K3 first at 33.4%. The third row is 0.8 points behind. The broader top-10 range is 6.6 points, so many of the published results sit in a relatively narrow band.

19 models have been evaluated on AA Tau3 Banking. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. AA Tau3 Banking is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About AA Tau3 Banking

Year

2026

Tasks

Banking tool-use workflows

Format

Task success rate

Difficulty

Agentic banking workflows

Stored separately from the broader Tau3-Bench lane.

BenchLM freshness & provenance

Version

AA Tau3 Banking 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does AA Tau3 Banking measure?

An independently evaluated Tau3 banking benchmark from Artificial Analysis.

Which model scores highest on AA Tau3 Banking?

Kimi K3 by Moonshot AI currently leads with a score of 33.4% on AA Tau3 Banking.

How many models are evaluated on AA Tau3 Banking?

19 AI models have been evaluated on AA Tau3 Banking on BenchLM.

Last updated: July 23, 2026 · BenchLM version AA Tau3 Banking 2026

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.