Model profile
o3-mini
Evidence coverage
10 of 323 tracked benchmarks are published. 5 are verified and 5 provisional. 5 of 8 categories are measured.
- Published / tracked
- 10 / 323
- Verified
- 5
- Provisional
- 5
- Categories with evidence
- 5 / 8
Evidence by category
- Agentic1 benchmarkReported
- Coding2 benchmarksMixed evidence
- Reasoning0 benchmarksNot measured
- Knowledge5 benchmarksMixed evidence
- Math1 benchmarkVerified
- Multilingual0 benchmarksNot measured
- Multimodal0 benchmarksNot measured
- Inst. Following1 benchmarkVerified
o3-mini ranks #136 out of 200 models on the public leaderboard with an overall score of 47.41/100. It also ranks #70 out of 99 on the verified leaderboard. While not a frontier model, it offers specific advantages depending on the use case.
o3-mini is a proprietary model with a 200K token context window. It uses explicit chain-of-thought reasoning, which typically improves performance on math and complex reasoning tasks at the cost of higher latency and token usage.
o3-mini sits inside the o3 family alongside o3, o3-pro. This profile currently has 10 of 323 tracked benchmarks. BenchLM only exposes non-generated benchmark rows publicly, so missing categories stay blank until a sourced evaluation is available.
Its strongest category is Instruction Following (#9), while its weakest is Coding (#75). This performance profile makes it a well-rounded choice across a range of tasks.
Peer position
Exact provisional scores and ranks for the closest listed peers. A score can appear before a model clears the evidence threshold for a rank, so equal scores can have different rank states.
Range 47.2–47.74
- Claude 3.5 SonnetAnthropicCompare#13247.74Claude 3.5 Sonnet is #132 with a score of 47.74.
- GLM-4.5-AirZ.AICompare#13347.7GLM-4.5-Air is #133 with a score of 47.7.
- Qwen3.5 FlashAlibabaCompare#13447.65Qwen3.5 Flash is #134 with a score of 47.65.
- Command A+CohereCompare#13547.51Command A+ is #135 with a score of 47.51.
- o3-miniCurrent modelOpenAI#13647.41o3-mini is #136 with a score of 47.41.
- Gemma 4 12BGoogleCompare#13747.29Gemma 4 12B is #137 with a score of 47.29.
- Qwen3.5 PlusAlibabaCompare#13847.2Qwen3.5 Plus is #138 with a score of 47.2.
Category percentile
More
Relative position among models eligible for each sourced category. A higher percentile means a stronger position within that category's ranked cohort; 100 is highest.
- Inst. Following73%Eligible cohort rank #9 of 31Category score 90.3
- Coding39%Eligible cohort rank #75 of 122Category score 48.3
Category evidence
Scores and ranks appear only where this model has published benchmark evidence. Categories without displayable source records remain not measured.
| Category | Score | Rank | Percentile | Weight | Benchmarks | Evidence |
|---|---|---|---|---|---|---|
| AgenticRank Not rankedWeight 22%1 benchmarkReported | 66.9 | Not ranked | Not available | 22% | 1 benchmark | Reported |
| CodingRank #75 of 122Percentile 39thWeight 20%2 benchmarksMixed sources | 48.3 | #75 of 122 | 39th | 20% | 2 benchmarks | Mixed sources |
| ReasoningWeight 17%0 benchmarksNot measured | Not measured | Not ranked | Not available | 17% | 0 benchmarks | Not measured |
| KnowledgeRank Not rankedWeight 12%5 benchmarksMixed sources | 65.9 | Not ranked | Not available | 12% | 5 benchmarks | Mixed sources |
| MathWeight 5%1 benchmarkVerified | Score pending | Not ranked | Not available | 5% | 1 benchmark | Verified |
| MultilingualWeight 7%0 benchmarksNot measured | Not measured | Not ranked | Not available | 7% | 0 benchmarks | Not measured |
| MultimodalWeight 12%0 benchmarksNot measured | Not measured | Not ranked | Not available | 12% | 0 benchmarks | Not measured |
| Inst. FollowingRank #9 of 31Percentile 73rdWeight 5%1 benchmarkVerified | 90.3 | #9 of 31 | 73rd | 5% | 1 benchmark | Verified |
Chatbot Arena performance
Scroll horizontally to inspect confidence intervals and vote counts.
| View | Elo | Confidence interval | Votes |
|---|---|---|---|
| Text Overall | 1348 | ±3.5 | 57,313 |
| Coding | 1416 | ±6.4 | 9,455 |
| Math | 1382 | ±8.5 | 4,719 |
| Instruction Following | 1344 | ±5.1 | 16,955 |
| Creative Writing | 1301 | ±7.0 | 8,182 |
| Multi-turn | 1340 | ±6.6 | 9,492 |
| Hard Prompts | 1370 | ±4.9 | 20,100 |
| Hard Prompts (English) | 1383 | ±6.0 | 11,286 |
| Longer Query | 1358 | ±6.5 | 9,224 |
Benchmark Details
Rows below have a displayable published verification record. Each source link and provenance note remains in the page HTML while its category is closed. Source-unverified manual rows and generated rows stay hidden.
Agentic1 benchmark
τ²-Bench Tool-Agent-User Evaluation
Coding2 benchmarks
Software Engineering Benchmark Verified
Artificial Analysis SciCode
Knowledge5 benchmarks
Graduate-Level Google-Proof Q&A
Massive Multitask Language Understanding
Artificial Analysis GPQA Diamond
Artificial Analysis Humanity's Last Exam
Math1 benchmark
American Invitational Mathematics Examination 2024
Inst. Following1 benchmark
Instruction-Following Eval
Frequently Asked Questions
How does o3-mini perform overall in AI benchmarks?
o3-mini has 10 published benchmark scores on BenchLM, but it does not yet have enough non-generated coverage to receive a global overall rank.
Is o3-mini good for knowledge and understanding?
o3-mini has visible benchmark coverage in knowledge and understanding, but BenchLM does not currently assign it a global category rank there.
Is o3-mini good for coding and programming?
o3-mini ranks #75 out of 122 models in coding and programming benchmarks with an average score of 48.3. There are stronger options in this category.
Is o3-mini good for mathematics?
o3-mini has visible benchmark coverage in mathematics, but BenchLM does not currently assign it a global category rank there.
Is o3-mini good for agentic tool use and computer tasks?
o3-mini has visible benchmark coverage in agentic tool use and computer tasks, but BenchLM does not currently assign it a global category rank there.
Is o3-mini good for instruction following?
o3-mini ranks #9 out of 31 models in instruction following benchmarks with an average score of 90.3. It is among the top performers in this category.
Which sibling models are related to o3-mini?
o3-mini belongs to the o3 family. Related variants on BenchLM include o3, o3-pro.
Does o3-mini have full benchmark coverage on BenchLM?
Not yet. o3-mini currently has 10 published benchmark scores out of the 323 benchmarks BenchLM tracks. BenchLM only exposes non-generated public benchmark rows, so missing categories stay blank until a sourced evaluation is available.
What is the context window size of o3-mini?
o3-mini has a published context window of 200K, which determines how much text it can process in a single interaction.
Related Resources
Choose with this week’s evidence
Join 2,000+ readers for ranking moves, new releases, pricing changes, and the evidence behind them.
Free. One email per week.