Benchmark profile
CharXiv Reasoning (CharXiv)
A scientific chart reasoning benchmark that tests whether models can understand, interpret, and reason about complex scientific visualizations including plots, diagrams, and data charts.
Data verifiedTop models on CharXiv — July 23, 2026
As of July 23, 2026, Claude Mythos 5 leads the CharXiv leaderboard with 93.5% , followed by Kimi K3 (91.3%) and Claude Opus 4.7 (Adaptive) (91%).
Claude Mythos 5
Anthropic
claude-mythos-5
Kimi K3
Moonshot AI
kimi-3
Claude Opus 4.7 (Adaptive)
Anthropic
claude-opus-4-7-max
Leaderboard (29 models)
ScoreAccording to BenchLM.ai, Claude Mythos 5 leads the CharXiv benchmark with a score of 93.5%, followed by Kimi K3 (91.3%) and Claude Opus 4.7 (Adaptive) (91%). The top models are clustered within 2.5 points, suggesting this benchmark is nearing saturation for frontier models.
29 models have been evaluated on CharXiv. The benchmark falls in the Multimodal & Grounded category. This category carries a 12% weight in BenchLM.ai's overall scoring system. Within that category, CharXiv contributes 25% of the category score, so strong performance here directly affects a model's overall ranking.
About CharXiv
Year
2024
Tasks
Scientific chart reasoning
Format
Chart understanding and reasoning
Difficulty
Scientific visualization reasoning
CharXiv evaluates a model's ability to reason about real-world scientific charts rather than simple visual QA. With-tools and without-tools variants isolate raw visual reasoning from tool-augmented performance.
BenchLM freshness & provenance
Version
CharXiv 2024
Refresh cadence
Annual
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does CharXiv measure?
A scientific chart reasoning benchmark that tests whether models can understand, interpret, and reason about complex scientific visualizations including plots, diagrams, and data charts.
Which model scores highest on CharXiv?
Claude Mythos 5 by Anthropic currently leads with a score of 93.5% on CharXiv.
How many models are evaluated on CharXiv?
29 AI models have been evaluated on CharXiv on BenchLM.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.