LLM Price vs Performance Chart
Find the most cost-effective AI model. Each dot is an LLM plotted by its provisional benchmark score (higher is better) against output token price (lower is better). Models on the efficiency frontier offer the best value at their price point.
Ministral 3 3B (Reasoning)
Score/$: 395.2 · $0.10/1M out
Claude Mythos 5
Score: 83.9 · $50.00/1M out
Ministral 3 3B (Reasoning)
Score: 39.5 · $0.10/1M out
107 priced models match these filters.
Top 10 Best Value Models (Overall)
Ranked by Score/$ ratio (benchmark score per dollar of output token cost)
| # | Model | Score | Output $/1M | Score/$ |
|---|---|---|---|---|
| 1 | Ministral 3 3B (Reasoning) Mistral | 39.5 | $0.10 | 395.2 |
| 2 | Ministral 3 8B (Reasoning) Mistral | 40.4 | $0.15 | 269.2 |
| 3 | Ministral 3 14B (Reasoning) Mistral | 49.3 | $0.20 | 246.7 |
| 4 | DeepSeek V4 Flash DeepSeek | 58.9 | $0.28 | 210.3 |
| 5 | DeepSeek V4 Flash (High) DeepSeek | 54 | $0.28 | 192.7 |
| 6 | Step 3.5 Flash StepFun | 55.1 | $0.30 | 183.7 |
| 7 | Ministral 3 3B Mistral | 18.3 | $0.10 | 182.7 |
| 8 | Ministral 3 14B Mistral | 34.5 | $0.20 | 172.3 |
| 9 | Ministral 3 8B Mistral | 21 | $0.15 | 139.7 |
| 10 | DeepSeek V3.2 DeepSeek | 55.4 | $0.42 | 131.9 |
Frequently Asked Questions
What is the LLM price-performance chart?
This chart plots each AI model by its benchmark score (vertical axis) against its API output price per million tokens (horizontal axis). Models in the upper-left quadrant offer the best value — high performance at low cost. The efficiency frontier line connects the best-value models at each price point.
What is the efficiency frontier?
The efficiency frontier (Pareto frontier) connects models where no other model offers both a higher score and a lower price. Models on this line represent the optimal price-performance tradeoff. If a model is below and to the right of the frontier, there exists a cheaper model with a better score.
Which LLM has the best price-to-performance ratio?
Currently, Ministral 3 3B (Reasoning) by Mistral offers the best overall value with a Score/$ ratio of 395.2. This means you get 395.2 benchmark points per dollar of output token cost.
How are scores calculated?
Overall scores shown in this chart use BenchLM's provisional ranking lane: a normalized weighted average across 8 benchmark categories, with bounded external calibration. The verified leaderboard is stricter and sourced-only, but this price-performance surface intentionally stays broader so value comparisons cover more models.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.