Best LLMs for Multimodal & Grounded — July 2026 Leaderboard
As of July 2026, the top multimodal & grounded model on the BenchLM leaderboard is Claude Opus 4.8 with a weighted multimodal & grounded score of 91.5.
Data refreshed:
Vision, document, and grounded enterprise workflow benchmarks
Decision lens: use provisional-ranked mode for broader public evidence and verified-ranked mode for source-only comparisons. A model can move between views as evidence coverage changes.
- Data refreshed
- July 23, 2026
- Provisional-ranked
- 29 of 290 models
- Verified-ranked
- 22 of 290 models
- Weighted evidence
- 2 of 24 benchmarks
24 tracked benchmarks
MMMU-Pro, OCRBench V2, olmOCR, VoxPopuli WER, OfficeQA Pro, MMMU-Pro w/ Python, OmniDocBench 1.5, Liquid Extract JSON Validity, Liquid Extract F1, Liquid Extract VLM Judge, GDPval-AA, Blueprint-Bench 2, MedXpertQA (MM), ZeroBench, Design2Code, Flame-VLM-Code, Vision2Web, ImageMining, MMSearch, MMSearch-Plus, SimpleVQA, Facts-VLM, V*, BabyVision
Scope: Vision, Document-office, GUI/web, Video
Evidence set: MMMU-Pro, OCRBench V2, olmOCR, VoxPopuli WER, OfficeQA Pro, MMMU-Pro w/ Python, OmniDocBench 1.5, Liquid Extract JSON Validity, Liquid Extract F1, Liquid Extract VLM Judge, GDPval-AA, Blueprint-Bench 2, MedXpertQA (MM), ZeroBench, Design2Code, Flame-VLM-Code, Vision2Web, ImageMining, MMSearch, MMSearch-Plus, SimpleVQA, Facts-VLM, V*, BabyVision
Scope: Vision, Document-office, GUI/web, Video
Best Multimodal & Grounded picks
BenchLM summaries for multimodal & grounded plus the practical tradeoffs users check next: open weights, price, speed, latency, and context.
Multimodal & Grounded Leaderboard
Primary score: weighted multimodal & grounded score. Higher values rank first. Use the Show metric control to change the value shown in each row.
Switch between provisional-ranked and verified-ranked modes to compare the broader public dataset with sourced-only rankings.
Filters
1 Claude Opus 4.8 Anthropic | 91.5% | 82 | — | — | — | — | 66.2% | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
2 Kimi K3 Moonshot AI | 90.3% | Est.77 | 81.6% | — | — | — | 63.3% | 83.4% | — | — | — | — | — | — | — | 23.0% | — | — | — | — | — | — | — | — | — | — |
3 GPT-5.6 Sol OpenAI | 84.3% | 80 | 83% | — | — | — | — | 84.6% | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
4 Gemini 3.5 Flash Google | 84% | 71 | 83.6% | — | — | — | — | — | — | — | — | — | — | 33.6% | — | — | — | — | — | — | — | — | — | — | — | — |
5 Gemini 3.1 Pro Google | 80.9% | 73 | 83.9% | — | — | — | 95%P | — | — | — | — | — | — | — | 81.3% | 29.0% | — | — | — | — | — | — | 72.4% | — | — | — |
6 Claude Sonnet 5 Anthropic | 78.9% | 79 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
7 Muse Spark Meta | 76.3% | 59 | 80.4% | — | — | — | — | — | — | — | — | — | — | — | 78.4% | 33.0% | — | — | — | — | — | — | 71.3% | — | — | — |
8 GPT-5.6 Terra OpenAI | 74.9% | 77 | 80.7% | — | — | — | — | 82% | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
9 Gemini 3 Pro Google | 73.2% | 58 | 81% | — | — | — | 92%P | — | 88.5%P | — | — | — | — | — | — | — | — | — | — | — | — | — | 73.2%P | — | 88.0% | — |
10 Qwen3.7 Plus Alibaba | 71.5% | 74 | 79% | 70.7% | — | — | — | — | 91.4% | — | — | — | — | — | 71.0% | — | — | — | — | — | — | 41.4% | 81.7% | — | — | — |
11 GPT-5.5 OpenAI | 70.1% | 74 | 81.2% | — | — | — | 54.1% | 83.2% | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
12 GPT-5.2 OpenAI | 69.3% | 55 | 79.5% | — | — | — | 95%P | — | 85.7%P | — | — | — | — | — | — | — | — | — | — | — | — | — | 55.8%P | — | 75.9% | — |
13 Grok 4.5 xAI | 69.1% | Est.72 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
14 GPT-5.4 OpenAI | 68.9% | 67 | 81.2% | — | — | — | 53.2% | 82.1% | — | — | — | — | — | — | 77.1% | 41.0% | — | — | — | — | — | — | 61.1% | — | — | — |
| 67.3% | 65 | 79.4% | — | — | — | — | 80.1% | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 96.9% | — | |
16 Qwen3.6 Plus Alibaba | 66.5% | 64 | 78.8% | — | — | — | — | — | 91.2%P | — | — | — | — | — | — | — | — | — | — | — | — | — | 67.3%P | — | 96.9% | — |
17 | 66.4% | 59 | 79% | — | — | — | 68%P | — | 90.8%P | — | — | — | — | — | — | — | — | — | — | — | — | — | 67.1%P | — | 95.8% | — |
18 GPT-5.6 Luna OpenAI | 65.9% | 73 | 78.4% | — | — | — | — | 79.5% | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
19 MiMo-V2.5 Xiaomi | 63.3% | 57 | 77.9% | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
20 | 54.2% | 53 | 75.8% | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 56.1% | — | 94.7% | — |
21 | 52.3% | 45 | 75.3% | — | — | — | — | — | 89.9% | — | — | — | — | — | — | — | — | — | — | — | — | — | 58.9% | — | — | — |
| 52% | 62 | 73.5% | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
23 Claude Opus 4.7 (Adaptive) Anthropic | 50.5% | 73 | — | — | — | — | 43.6% | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
24 | 48.2% | 61 | 78.1% | — | — | — | 45.1% | — | 91.6% | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
25 Grok 4.20 xAI | 34.9% | 47 | 75.2% | — | — | — | — | — | — | — | — | — | — | — | 65.8% | — | — | — | — | — | — | — | 57.4% | — | — | — |
Top AI Models for Multimodal & Grounded — July 2026
As of July 2026, Claude Opus 4.8 leads the provisional multimodal & grounded leaderboard with a score of 91.5%, followed by Kimi K3 (90.3%) and GPT-5.6 Sol (84.3%). BenchLM is currently showing 29 provisional-ranked models and 22 verified-ranked models in this category.
What changed
Claude Mythos Preview leads multimodal with the strongest MMMU-Pro score.
GPT-5.4 close behind with strong OfficeQA Pro and MMMU-Pro results.
Claude Opus 4.7 adds official CharXiv visual reasoning coverage.
Top models by benchmark
Frontier multimodal reasoning benchmark spanning charts, diagrams, tables, and visual academic problems(45% of category score)
Score in Context
What these scores mean
Multimodal & Grounded carries a 12% weight in overall scoring. The weighted score blends MMMU-Pro (academic multimodal reasoning), OfficeQA Pro (enterprise document understanding), and CharXiv (visual chart and figure reasoning). A model can know facts in text and still fail when the information is in a chart, screenshot, or spreadsheet — this category measures that gap.
Known limitations
Not all models support image input — text-only models are excluded from this category entirely. OfficeQA Pro and CharXiv coverage is still building, so rankings should be read as a blend of available public evidence rather than a complete visual capability profile. Enterprise-specific document formats (scanned PDFs, handwritten notes) remain under-tested by all benchmarks.
How we weight
Multimodal & Grounded carries a 12% weight in BenchLM.ai's overall scoring. It remains important for enterprise copilots and document-heavy workflows where models need to interpret visuals, screenshots, and scanned artifacts.
This category tests whether a model can read the world as it actually appears in products: screenshots, charts, scanned documents, and mixed visual-text artifacts. See the multimodal leaderboard or compare with knowledge benchmarks.
Leaderboards exclude benchmark rows that BenchLM generated from other scores or cloned from reference models. When a weighted benchmark is missing after that filter, the category falls back to the remaining trustworthy public rows instead of filling the gap with synthetic values.
The full scoring rules, freshness handling, and runtime/pricing caveats live on the BenchLM methodology page.
Scroll horizontally to read the full evidence ledger.
| Benchmark | Weight | Status | Description |
|---|---|---|---|
| MMMU-Pro | 45% | Weighted | Frontier multimodal reasoning benchmark spanning charts, diagrams, tables, and visual academic problems |
| OCRBench V2 | — | Display only | A native OCR benchmark for reading text from images across multilingual scripts, low-quality scans, handwriting, structured layouts, charts, and screenshots. |
| olmOCR | — | Display only | An end-to-end document understanding benchmark over long, layout-rich PDFs with tables, equations, headers, footnotes, and multi-column flows. |
| VoxPopuli WER | — | Display only | A speech-recognition benchmark on the cleaned Artificial Analysis VoxPopuli subset, reported as word error rate where lower is better. |
| OfficeQA Pro | 30% | Weighted | Grounded office and enterprise document benchmark |
| MMMU-Pro w/ Python | — | Display only | Tool-augmented MMMU-Pro variant that allows Python assistance during multimodal reasoning |
| OmniDocBench 1.5 | — | Display only | Document understanding benchmark measured by edit distance on complex document extraction tasks |
| Liquid Extract JSON Validity | — | Display only | A display-only Liquid AI extraction metric measuring strict JSON parseability. |
| Liquid Extract F1 | — | Display only | A display-only Liquid AI extraction metric measuring requested-field agreement. |
| Liquid Extract VLM Judge | — | Display only | A display-only Liquid AI extraction metric measuring judged agreement with the source image. |
| GDPval-AA | — | Display only | An evaluation focused on professional domain expertise and task delivery quality in office-style knowledge work. |
| Blueprint-Bench 2 | — | Display only | Agentic spatial reasoning benchmark reported as a normalized score. |
| MedXpertQA (MM) | — | Display only | A clinically grounded multimodal medical multiple-choice benchmark with image inputs. |
| ZeroBench | — | Display only | A multi-step visual reasoning benchmark with pass@5 reporting and optional tool use. |
| Design2Code | — | Display only | Multimodal coding benchmark for turning visual designs into working frontend implementations. |
| Flame-VLM-Code | — | Display only | Vision-language coding benchmark for generating correct code from visual and multimodal inputs. |
| Vision2Web | — | Display only | Benchmark for converting visual references into functional web implementations. |
| ImageMining | — | Display only | Multimodal retrieval and extraction benchmark over image-heavy task settings. |
| MMSearch | — | Display only | Multimodal search benchmark for retrieval and grounded answering across mixed-media inputs. |
| MMSearch-Plus | — | Display only | A harder MMSearch variant for multimodal retrieval and grounded tool-use workflows. |
| SimpleVQA | — | Display only | Visual question answering benchmark focused on straightforward image-grounded understanding. |
| Facts-VLM | — | Display only | Grounded multimodal factuality benchmark for evidence-linked answer correctness. |
| V* | — | Display only | Vision-centric benchmark for high-level multimodal reasoning and perception quality. |
| BabyVision | — | Display only | A multimodal benchmark for fine-grained visual perception and grounded reasoning tasks. |
About Multimodal & Grounded Benchmarks
Frontier multimodal reasoning benchmark spanning charts, diagrams, tables, and visual academic problems
Common questions
What do multimodal and grounded LLM benchmarks measure?
They measure whether models can reason over images, charts, documents, and office artifacts instead of only plain-text prompts.
Which benchmarks matter most here?
MMMU-Pro is a strong frontier multimodal reasoning benchmark, OfficeQA Pro is useful for grounded document and office-style enterprise workflows, and CharXiv tests chart and figure reasoning.
Why is this category separate from knowledge?
A model can know facts in text and still be weak when the information is embedded in a chart, screenshot, PDF, or spreadsheet. This category measures that gap directly.
Multimodal benchmark updates
Multimodal rankings are heating up. Get the weekly update.
One email each week. Unsubscribe anytime.