Benchmark profile
CharXiv Reasoning without tools (CharXiv w/o tools)
Tool-free variant of CharXiv that isolates raw visual reasoning ability without code execution or tool augmentation.
Data verifiedBenchmark score on CharXiv w/o tools — July 23, 2026
BenchLM mirrors the published score view for CharXiv w/o tools. Claude Mythos 5 leads the public snapshot at 88.9% , followed by Kimi K3 (84.8%) and Claude Opus 4.7 (Adaptive) (82.1%). BenchLM does not use these results to rank models overall.
Claude Mythos 5
Anthropic
claude-mythos-5
Kimi K3
Moonshot AI
kimi-3
Claude Opus 4.7 (Adaptive)
Anthropic
claude-opus-4-7-max
Benchmark score table (6 models)
ScoreThe published CharXiv w/o tools snapshot places Claude Mythos 5 first at 88.9%. The third row is 6.8 points behind. The broader top-10 range is 11.9 points, so the table still separates the published systems.
6 models have been evaluated on CharXiv w/o tools. The benchmark falls in the Multimodal & Grounded category. This category carries a 12% weight in BenchLM.ai's overall scoring system. CharXiv w/o tools is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About CharXiv w/o tools
Year
2024
Tasks
Scientific chart reasoning (tool-free)
Format
Chart understanding without tools
Difficulty
Scientific visualization reasoning
The tool-free CharXiv variant measures pure multimodal reasoning. Mythos Preview scores 86.1% without tools vs 93.2% with tools, demonstrating strong baseline chart reasoning.
BenchLM freshness & provenance
Version
CharXiv w/o tools 2024
Refresh cadence
Annual
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does CharXiv w/o tools measure?
Tool-free variant of CharXiv that isolates raw visual reasoning ability without code execution or tool augmentation.
Which model scores highest on CharXiv w/o tools?
Claude Mythos 5 by Anthropic currently leads with a score of 88.9% on CharXiv w/o tools.
How many models are evaluated on CharXiv w/o tools?
6 AI models have been evaluated on CharXiv w/o tools on BenchLM.
Compare Top Models on CharXiv w/o tools
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.