Skip to main content

BenchLM recommendation

Best Multimodal LLMs in 2026

Data verified

As of July 23, 2026, the top model in best multimodal llms on the BenchLM leaderboard is Claude Opus 4.8 with a score of 91.5.

Last verified: July 23, 2026

This page ranks models by BenchLM's multimodal & grounded category: image understanding, document and chart reading, and visually grounded reasoning. Note the scope: these scores measure how well a model understands images — not how well it generates them. Image-generation systems are a different product class and are not ranked here.

Unless noted otherwise, ranking surfaces on this page use BenchLM's provisional leaderboard lane rather than the stricter sourced-only verified leaderboard.

Bottom line: Gemini 3 Pro Deep Think leads multimodal understanding at 95, with Claude Mythos 5 and GPT-5.1 behind it. For image generation, look at dedicated image models — this table measures visual understanding.

Claude Opus 4.8 leads this ranking with a score of 91.5, followed by Kimi K3 (90.3) and GPT-5.6 Sol (84.3). There is meaningful separation between the top models, suggesting genuine performance differences.

The best open-weight option is Kimi K2.6 (ranked #15 with a score of 67.3). Proprietary models hold a clear advantage in this category, though open-weight options may suffice for less demanding use cases.

This ranking is based on provisional weighted averages across the scoring benchmarks in multimodalGrounded tracked by BenchLM.ai. For detailed model profiles, click any model name below. To compare two specific models head-to-head, use the "vs #" links.

What changed

Gemini 3 Pro Deep Think leads multimodal & grounded understanding at 95.

Claude Mythos 5 strong #2 at 92.8 on grounded visual reasoning.

GPT-5.1 still a top-3 multimodal row at 91.8 — and cheap at $1.25/$10.

How to choose

Full Rankings (27 models)

1
Claude Opus 4.8
Anthropic·Proprietary·1M

91.5

prov. avg

2
Kimi K3
Moonshot AI·Pending·1.05M

90.3

prov. avg

3
GPT-5.6 Sol
OpenAI·Proprietary·1M

84.3

prov. avg

4
Gemini 3.5 Flash
Google·Proprietary·1M

84

prov. avg

5
Gemini 3.1 Pro
Google·Proprietary·1M

80.9

prov. avg

6
Claude Sonnet 5
Anthropic·Proprietary·1M

78.9

prov. avg

7
Muse Spark
Meta·Proprietary·262K

76.3

prov. avg

8
GPT-5.6 Terra
OpenAI·Proprietary·1M

74.9

prov. avg

9
Gemini 3 Pro
Google·Proprietary·2M

73.2

prov. avg

10
Qwen3.7 Plus
Alibaba·Proprietary·1M

71.5

prov. avg

11
GPT-5.5
OpenAI·Proprietary·1M

70.1

prov. avg

12
GPT-5.2
OpenAI·Proprietary·400K

69.3

prov. avg

13
Grok 4.5
xAI·Proprietary·500K

69.1

prov. avg

14
GPT-5.4
OpenAI·Proprietary·1.05M

68.9

prov. avg

15
Kimi K2.6
Moonshot AI·Open Weight·256K

67.3

prov. avg

16
Qwen3.6 Plus
Alibaba·Proprietary·1M

66.5

prov. avg

17
Qwen3.5 397B
Alibaba·Open Weight·128K

66.4

prov. avg

18
GPT-5.6 Luna
OpenAI·Proprietary·1M

65.9

prov. avg

19
MiMo-V2.5
Xiaomi·Proprietary·1M

63.3

prov. avg

20
Qwen3.6-27B
Alibaba·Open Weight·262K

54.2

prov. avg

21
Qwen3.6-35B-A3B
Alibaba·Open Weight·262K

52.3

prov. avg

22
Inkling
Thinking Machines Lab·Open Weight·1M

52

prov. avg

23
Claude Opus 4.7 (Adaptive)
Anthropic·Proprietary·1M

50.5

prov. avg

24
MiniMax M3
MiniMax·Open Weight·1M

48.2

prov. avg

25
Grok 4.20
xAI·Proprietary·2M

34.9

prov. avg

26
Claude Opus 4.5
Anthropic·Proprietary·200K

28.4

prov. avg

27
Command A+
Cohere·Open Weight·128K

7

prov. avg

Key Takeaways

The top model is Claude Opus 4.8 by Anthropic with a provisional score of 91.5.

The best open-weight model is Kimi K2.6 at position #15.

27 models are included in this ranking.

Score in Context

What these scores mean

The multimodal & grounded score blends image understanding, document/chart reading, and visually grounded reasoning benchmarks. Higher means the model more reliably extracts and reasons over visual content.

Known limitations

These are understanding scores — they say nothing about image generation quality. OCR-heavy production workloads should also test resolution limits and page-count behavior, which benchmarks only partially cover.

Best Multimodal LLMs FAQ

What is the best multimodal LLM?

Gemini 3 Pro Deep Think currently leads BenchLM's multimodal & grounded category at 95, ahead of Claude Mythos 5 (92.8) and GPT-5.1 (91.8). The table above recomputes with every data refresh. For most teams GPT-5.1 is the value pick — top-3 understanding at $1.25/$10 per million tokens.

What is the best LLM for image generation?

This page ranks image understanding, not generation — LLM leaderboards measure how well models read and reason over images. Image generation is a separate product class (diffusion and autoregressive image models) that BenchLM's categories do not currently score. Treat any "LLM" ranking of image generation with skepticism.

Which LLM is best for reading documents and PDFs?

Document understanding tracks the multimodal & grounded category closely: Gemini 3 Pro Deep Think leads, and long-context support matters as much as vision quality for multi-hundred-page work — see the large-context rankings for models pairing 1M-token windows with strong vision.

Can these models understand charts and tables?

The top rows handle standard charts and tables well — chart-reading benchmarks feed this category. Reliability drops on dense dashboards, small text, and unusual chart types, so test your hardest real documents rather than trusting a single aggregate score.

Last updated: July 23, 2026

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.