Skip to main content

Model profile

Grok 4.3

xAISupersededReleased Apr 30, 2026
Data verified
Superseded:xAI has released newer models in this line —Grok 4.5
Overall Score
65.1Public #31 of 200Verified #26 of 99
Arena Elo
1443
Eligible category ranks
4of 8
Price (1M tokens)
$1.25 in / $2.5 out
API pricing
Speed
209tok/s
Context
1M

Evidence coverage

25 of 323 tracked benchmarks are published. 2 are verified and 23 provisional. 6 of 8 categories are measured.

Updated July 23, 2026Methodology
Published / tracked
25 / 323
Verified
2
Provisional
23
Categories with evidence
6 / 8

Evidence by category

  • Agentic7 benchmarks
    Mixed evidence
  • Coding3 benchmarks
    Reported
  • Reasoning2 benchmarks
    Reported
  • Knowledge8 benchmarks
    Reported
  • Math0 benchmarks
    Not measured
  • Multilingual0 benchmarks
    Not measured
  • Multimodal3 benchmarks
    Reported
  • Inst. Following2 benchmarks
    Reported
ProprietaryReasoning
Confidence:
Low
base

Grok 4.3 ranks #31 out of 200 models on the public leaderboard with an overall score of 65.1/100. It also ranks #26 out of 99 on the verified leaderboard. While not a frontier model, it offers specific advantages depending on the use case.

Grok 4.3 is a proprietary model with a 1M token context window. It uses explicit chain-of-thought reasoning, which typically improves performance on math and complex reasoning tasks at the cost of higher latency and token usage.

This profile currently has 25 of 323 tracked benchmarks. BenchLM only exposes non-generated benchmark rows publicly, so missing categories stay blank until a sourced evaluation is available.

Its strongest category is Instruction Following (#6), while its weakest is Agentic (#119). This performance profile makes it a well-rounded choice across a range of tasks.

Peer position

Exact provisional scores and ranks for the closest listed peers. A score can appear before a model clears the evidence threshold for a rank, so equal scores can have different rank states.

Range 64.1865.32

  1. Claude Sonnet 5
    Anthropic
    #2965.32
    Claude Sonnet 5 is #29 with a score of 65.32.
    Compare
  2. Qwen3.6 Plus
    Alibaba
    #3065.2
    Qwen3.6 Plus is #30 with a score of 65.2.
    Compare
  3. Grok 4.3Current model
    xAI
    #3165.1
    Grok 4.3 is #31 with a score of 65.1.
  4. Claude Sonnet 4.6
    Anthropic
    #3265.07
    Claude Sonnet 4.6 is #32 with a score of 65.07.
    Compare
  5. Gemini 3.5 Flash
    Google
    #3364.75
    Gemini 3.5 Flash is #33 with a score of 64.75.
    Compare
  6. Claude Opus 4.5
    Anthropic
    #3464.22
    Claude Opus 4.5 is #34 with a score of 64.22.
    Compare
  7. Claude Opus 4.6 (Adaptive)
    Anthropic
    #3564.18
    Claude Opus 4.6 (Adaptive) is #35 with a score of 64.18.
    Compare

Category percentile

More

Relative position among models eligible for each sourced category. A higher percentile means a stronger position within that category's ranked cohort; 100 is highest.

  1. Inst. Following83%
    Eligible cohort rank #6 of 31Category score 91.2
  2. Knowledge16%
    Eligible cohort rank #44 of 52Category score 57.0
  3. Coding12%
    Eligible cohort rank #107 of 122Category score 39.4
  4. Agentic0%
    Eligible cohort rank #119 of 119Category score 6.3

Category evidence

Scores and ranks appear only where this model has published benchmark evidence. Categories without displayable source records remain not measured.

Category scores, ranks, weighting, benchmark coverage, and evidence status
CategoryScore
AgenticRank #119 of 119Percentile 0thWeight 22%7 benchmarksMixed sources6.3
CodingRank #107 of 122Percentile 12thWeight 20%3 benchmarksReported39.4
ReasoningWeight 17%2 benchmarksReportedScore pending
KnowledgeRank #44 of 52Percentile 16thWeight 12%8 benchmarksReported57.0
MathWeight 5%0 benchmarksNot measuredNot measured
MultilingualWeight 7%0 benchmarksNot measuredNot measured
MultimodalRank Not rankedWeight 12%3 benchmarksReported64.3
Inst. FollowingRank #6 of 31Percentile 83rdWeight 5%2 benchmarksReported91.2

Chatbot Arena performance

Scroll horizontally to inspect confidence intervals and vote counts.

Chatbot Arena Elo, confidence interval, and vote count by evaluation view
ViewEloConfidence intervalVotes
Text Overall1443±4.345,525
Coding1487±6.812,913
Math1418±13.12,259
Instruction Following1413±6.415,451
Creative Writing1428±8.17,682
Multi-turn1451±7.88,348
Hard Prompts1458±5.230,172
Hard Prompts (English)1461±6.515,126
Longer Query1445±6.120,198

Benchmark Details

Rows below have a displayable published verification record. Each source link and provenance note remains in the page HTML while its category is closed. Source-unverified manual rows and generated rows stay hidden.

Agentic7 benchmarks
τ²-bench resultsSecondary exact

τ²-Bench Tool-Agent-User Evaluation

97.7%Display only
Source: Artificial Analysis: Grok 4.3Provenance: Secondary exact
GDPval-AASecondary exact

GDPval-AA normalized

29.2%Display only
Source: OpenRouter: Grok 4.3 benchmarksProvenance: OpenRouter renders the Artificial Analysis agentic card reporting GDPval-AA as a normalized 49.9% score, separate from Elo-style GDPval-AA rows.
AA Agentic IndexReported

Artificial Analysis Agentic Index

24.1%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
APEX-Agents-AAReported
17.0%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
GDPval-AAReported
1085Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Gert LabsBenchmark exact

Gert Labs Composite Game Benchmark

43.86%Display only
Source: Gert Labs rankingsProvenance: Gert Labs reports this composite leaderboard score in the public rankings API. BenchLM scales the source gscore from 0-1 to 0-100 and stores it as a display-only agentic benchmark.
ResearchClawBenchBenchmark exact
12.4%Display only
Source: ResearchClawBench leaderboardProvenance: ResearchClawBench reports this model as ResearchHarness (Grok-4.3) in the official Pass@1 leaderboard. BenchLM stores the one-decimal RADS average on the local ResearchClawBench display key and excludes it from weighted rankings.
Coding3 benchmarks
SciCodeSecondary exact

Scientific Code Benchmark

47.3%Weighted 16%
Source: Artificial Analysis: Grok 4.3Provenance: Secondary exact
AA Coding IndexReported

Artificial Analysis Coding Index

42.3%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-SciCodeReported

Artificial Analysis SciCode

47.3%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Reasoning2 benchmarks
AA-LCRSecondary exact

Artificial Analysis Long Context Reasoning

64.3%Display only
Source: OpenRouter: Grok 4.3 benchmarksProvenance: OpenRouter renders the Artificial Analysis reasoning card reporting AA-LCR at 64.3%.
CritPtSecondary exact

Critical Physics Tasks

8.0%Display only
Source: OpenRouter: Grok 4.3 benchmarksProvenance: OpenRouter renders the Artificial Analysis reasoning card reporting CritPt at 8.0%.
Knowledge8 benchmarks
HLESecondary exact

Humanity's Last Exam

35%Weighted 45%
Source: Artificial Analysis: Grok 4.3Provenance: Secondary exact
GPQASecondary exact

Graduate-Level Google-Proof Q&A

90.1%Weighted 7%
Source: Artificial Analysis: Grok 4.3Provenance: Artificial Analysis reports GPQA at 90.1%. BenchLM maps that onto the site GPQA row for comparability.
Artificial Analysis Intelligence IndexSecondary exact
37.6%Display only
Source: Artificial Analysis: Grok 4.3Provenance: Artificial Analysis reports Grok 4.3 at 53.2 on the Intelligence Index.
AA-Omniscience AccuracySecondary exact

Artificial Analysis Omniscience Accuracy

34.6%Display only
Source: OpenRouter: Grok 4.3 benchmarksProvenance: OpenRouter renders the Artificial Analysis knowledge card reporting AA-Omniscience Accuracy at 34.6%.
AA-Omniscience Hallucination RateSecondary exact

Artificial Analysis Omniscience Hallucination Rate

25.0%Display only
Source: OpenRouter: Grok 4.3 benchmarksProvenance: OpenRouter renders the Artificial Analysis knowledge card reporting AA-Omniscience Hallucination Rate at 75.0%. Lower is better for this row.
AA-GPQA DiamondReported

Artificial Analysis GPQA Diamond

90.1%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-HLEReported

Artificial Analysis Humanity's Last Exam

35.0%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-Omniscience IndexReported

Artificial Analysis Omniscience Index

18.3%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Multimodal3 benchmarks
MMMU-ProSecondary exact

Massive Multi-discipline Multimodal Understanding Pro

78.1%Weighted 45%
Source: Artificial Analysis: Grok 4.3Provenance: Secondary exact
Design Arena WebsiteSecondary exact

Design Arena Website Elo

1225Display only
Source: OpenRouter: Grok 4.3 benchmarksProvenance: OpenRouter renders the Design Arena website card reporting 1294 Elo, 56.5% win rate, 166.3s average generation time, and Top 13%. BenchLM stores the Elo as the display value.
AA-MMMU-ProReported

Artificial Analysis MMMU-Pro

78.1%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Inst. Following2 benchmarks
IFBenchSecondary exact

Instruction Following Benchmark

81.3%Weighted 65%
Source: Artificial Analysis: Grok 4.3Provenance: Secondary exact
AA-IFBenchReported

Artificial Analysis IFBench

81.3%Display only
Source: Artificial Analysis: ifbench leaderboardProvenance: Display-only row synced from the current Artificial Analysis evaluation leaderboard. It is excluded from BenchLM weighted scoring.

Frequently Asked Questions

How does Grok 4.3 perform overall in AI benchmarks?

Grok 4.3 has 25 published benchmark scores on BenchLM, but it does not yet have enough non-generated coverage to receive a global overall rank.

Is Grok 4.3 good for knowledge and understanding?

Grok 4.3 ranks #44 out of 52 models in knowledge and understanding benchmarks with an average score of 57. There are stronger options in this category.

Is Grok 4.3 good for coding and programming?

Grok 4.3 ranks #107 out of 122 models in coding and programming benchmarks with an average score of 39.4. There are stronger options in this category.

Is Grok 4.3 good for reasoning and logic?

Grok 4.3 has visible benchmark coverage in reasoning and logic, but BenchLM does not currently assign it a global category rank there.

Is Grok 4.3 good for agentic tool use and computer tasks?

Grok 4.3 ranks #119 out of 119 models in agentic tool use and computer tasks benchmarks with an average score of 6.3. There are stronger options in this category.

Is Grok 4.3 good for multimodal and grounded tasks?

Grok 4.3 has visible benchmark coverage in multimodal and grounded tasks, but BenchLM does not currently assign it a global category rank there.

Is Grok 4.3 good for instruction following?

Grok 4.3 ranks #6 out of 31 models in instruction following benchmarks with an average score of 91.2. It is among the top performers in this category.

Does Grok 4.3 have full benchmark coverage on BenchLM?

Not yet. Grok 4.3 currently has 25 published benchmark scores out of the 323 benchmarks BenchLM tracks. BenchLM only exposes non-generated public benchmark rows, so missing categories stay blank until a sourced evaluation is available.

What is the context window size of Grok 4.3?

Grok 4.3 has a published context window of 1M, which determines how much text it can process in a single interaction.

Last updated: July 23, 2026 · Runtime metrics stay blank until BenchLM has a sourced snapshot.

Choose with this week’s evidence

Join 2,000+ readers for ranking moves, new releases, pricing changes, and the evidence behind them.

Free. One email per week.