Model profile
GPT-4o mini
Evidence coverage
10 of 323 tracked benchmarks are published. 0 are verified and 10 provisional. 5 of 8 categories are measured.
- Published / tracked
- 10 / 323
- Verified
- 0
- Provisional
- 10
- Categories with evidence
- 5 / 8
Evidence by category
- Agentic3 benchmarksReported
- Coding2 benchmarksReported
- Reasoning0 benchmarksNot measured
- Knowledge3 benchmarksReported
- Math0 benchmarksNot measured
- Multilingual0 benchmarksNot measured
- Multimodal1 benchmarkReported
- Inst. Following1 benchmarkReported
GPT-4o mini ranks #183 out of 200 models on the public leaderboard with an overall score of 37.87/100. It also ranks #82 out of 99 on the verified leaderboard. While not a frontier model, it offers specific advantages depending on the use case.
GPT-4o mini is a proprietary model with a 128K token context window. It processes queries without explicit chain-of-thought reasoning, offering faster response times and lower token usage.
GPT-4o mini sits inside the GPT-4o family alongside GPT-4o. This profile currently has 10 of 323 tracked benchmarks. BenchLM only exposes non-generated benchmark rows publicly, so missing categories stay blank until a sourced evaluation is available.
Its strongest category is Agentic (#110), while its weakest is Coding (#112). This performance profile makes it particularly useful for coding agents, browser research, and computer-use workflows.
Peer position
Exact provisional scores and ranks for the closest listed peers. A score can appear before a model clears the evidence threshold for a rank, so equal scores can have different rank states.
Range 36.55–39.52
- Ministral 3 3B (Reasoning)MistralCompare#17739.52Ministral 3 3B (Reasoning) is #177 with a score of 39.52.
- Mistral 8x7B v0.2MistralCompare#17839.13Mistral 8x7B v0.2 is #178 with a score of 39.13.
- Grok Code Fast 1xAICompare#18038.64Grok Code Fast 1 is #180 with a score of 38.64.
- Granite-4.0-350MIBMCompare#18138.32Granite-4.0-350M is #181 with a score of 38.32.
- Granite-4.0-H-350MIBMCompare#18238.32Granite-4.0-H-350M is #182 with a score of 38.32.
- GPT-4o miniCurrent modelOpenAI#18337.87GPT-4o mini is #183 with a score of 37.87.
- Claude 4.1 Opus ThinkingAnthropicCompare#18436.55Claude 4.1 Opus Thinking is #184 with a score of 36.55.
Category percentile
More
Relative position among models eligible for each sourced category. A higher percentile means a stronger position within that category's ranked cohort; 100 is highest.
- Agentic8%Eligible cohort rank #110 of 119Category score 33.7
- Coding8%Eligible cohort rank #112 of 122Category score 35.6
Category evidence
Scores and ranks appear only where this model has published benchmark evidence. Categories without displayable source records remain not measured.
| Category | Score | Rank | Percentile | Weight | Benchmarks | Evidence |
|---|---|---|---|---|---|---|
| AgenticRank #110 of 119Percentile 8thWeight 22%3 benchmarksReported | 33.7 | #110 of 119 | 8th | 22% | 3 benchmarks | Reported |
| CodingRank #112 of 122Percentile 8thWeight 20%2 benchmarksReported | 35.6 | #112 of 122 | 8th | 20% | 2 benchmarks | Reported |
| ReasoningWeight 17%0 benchmarksNot measured | Not measured | Not ranked | Not available | 17% | 0 benchmarks | Not measured |
| KnowledgeWeight 12%3 benchmarksReported | Score pending | Not ranked | Not available | 12% | 3 benchmarks | Reported |
| MathWeight 5%0 benchmarksNot measured | Not measured | Not ranked | Not available | 5% | 0 benchmarks | Not measured |
| MultilingualWeight 7%0 benchmarksNot measured | Not measured | Not ranked | Not available | 7% | 0 benchmarks | Not measured |
| MultimodalRank Not rankedWeight 12%1 benchmarkReported | 60.8 | Not ranked | Not available | 12% | 1 benchmark | Reported |
| Inst. FollowingWeight 5%1 benchmarkReported | Score pending | Not ranked | Not available | 5% | 1 benchmark | Reported |
Chatbot Arena performance
Scroll horizontally to inspect confidence intervals and vote counts.
| View | Elo | Confidence interval | Votes |
|---|---|---|---|
| Text Overall | 1317 | Not available | Not available |
Benchmark Details
Rows below have a displayable published verification record. Each source link and provenance note remains in the page HTML while its category is closed. Source-unverified manual rows and generated rows stay hidden.
Agentic3 benchmarks
Artificial Analysis Agentic Index
GDPval-AA normalized
Coding2 benchmarks
Artificial Analysis SciCode
Artificial Analysis Coding Index
Knowledge3 benchmarks
Artificial Analysis GPQA Diamond
Artificial Analysis Humanity's Last Exam
Multimodal1 benchmark
Artificial Analysis MMMU-Pro
Inst. Following1 benchmark
Artificial Analysis IFBench
Frequently Asked Questions
How does GPT-4o mini perform overall in AI benchmarks?
GPT-4o mini has 10 published benchmark scores on BenchLM, but it does not yet have enough non-generated coverage to receive a global overall rank.
Is GPT-4o mini good for knowledge and understanding?
GPT-4o mini has visible benchmark coverage in knowledge and understanding, but BenchLM does not currently assign it a global category rank there.
Is GPT-4o mini good for coding and programming?
GPT-4o mini ranks #112 out of 122 models in coding and programming benchmarks with an average score of 35.6. There are stronger options in this category.
Is GPT-4o mini good for agentic tool use and computer tasks?
GPT-4o mini ranks #110 out of 119 models in agentic tool use and computer tasks benchmarks with an average score of 33.7. There are stronger options in this category.
Is GPT-4o mini good for multimodal and grounded tasks?
GPT-4o mini has visible benchmark coverage in multimodal and grounded tasks, but BenchLM does not currently assign it a global category rank there.
Is GPT-4o mini good for instruction following?
GPT-4o mini has visible benchmark coverage in instruction following, but BenchLM does not currently assign it a global category rank there.
Which sibling models are related to GPT-4o mini?
GPT-4o mini belongs to the GPT-4o family. Related variants on BenchLM include GPT-4o.
Does GPT-4o mini have full benchmark coverage on BenchLM?
Not yet. GPT-4o mini currently has 10 published benchmark scores out of the 323 benchmarks BenchLM tracks. BenchLM only exposes non-generated public benchmark rows, so missing categories stay blank until a sourced evaluation is available.
What is the context window size of GPT-4o mini?
GPT-4o mini has a published context window of 128K, which determines how much text it can process in a single interaction.
Related Resources
Choose with this week’s evidence
Join 2,000+ readers for ranking moves, new releases, pricing changes, and the evidence behind them.
Free. One email per week.