Model profile
Mellum2-12B-A2.5B-Thinking
Evidence coverage
5 of 323 tracked benchmarks are published. 5 are verified and 0 provisional. 3 of 8 categories are measured.
- Published / tracked
- 5 / 323
- Verified
- 5
- Provisional
- 0
- Categories with evidence
- 3 / 8
Evidence by category
- Agentic1 benchmarkVerified
- Coding0 benchmarksNot measured
- Reasoning0 benchmarksNot measured
- Knowledge3 benchmarksVerified
- Math0 benchmarksNot measured
- Multilingual0 benchmarksNot measured
- Multimodal0 benchmarksNot measured
- Inst. Following1 benchmarkVerified
BenchLM is tracking Mellum2-12B-A2.5B-Thinking, but this profile is currently excluded from the public leaderboard because it still lacks enough non-generated benchmark coverage to rank safely. Only non-generated public benchmark rows appear below.
Mellum2-12B-A2.5B-Thinking is a open weight model with a 128K token context window. It uses explicit chain-of-thought reasoning, which typically improves performance on math and complex reasoning tasks at the cost of higher latency and token usage.
Mellum2-12B-A2.5B-Thinking sits inside the Mellum2 12B-A2.5B family alongside Mellum2-12B-A2.5B-Instruct. This profile currently has 5 of 323 tracked benchmarks. BenchLM only exposes non-generated benchmark rows publicly, so missing categories stay blank until a sourced evaluation is available.
Its strongest category is Instruction Following (#27), while its weakest is Agentic (#62). This performance profile makes it a well-rounded choice across a range of tasks.
Peer position
Exact provisional scores and ranks for the closest listed peers. A score can appear before a model clears the evidence threshold for a rank, so equal scores can have different rank states.
Range 77.44–83.93
- Claude Mythos 5AnthropicCompare#183.93Claude Mythos 5 is #1 with a score of 83.93.
- Claude Fable 5AnthropicCompare#283.68Claude Fable 5 is #2 with a score of 83.68.
- GPT-5.6 SolOpenAICompare#381.96GPT-5.6 Sol is #3 with a score of 81.96.
- Kimi K3Moonshot AICompare#480.96Kimi K3 is #4 with a score of 80.96.
- Claude Opus 4.8AnthropicCompare#578.34Claude Opus 4.8 is #5 with a score of 78.34.
- Muse Spark 1.1MetaCompare#677.44Muse Spark 1.1 is #6 with a score of 77.44.
- Mellum2-12B-A2.5B-ThinkingCurrent modelJetBrainsUnrankedNot measuredMellum2-12B-A2.5B-Thinking is Unranked with a score of Not measured.
Category percentile
More
Relative position among models eligible for each sourced category. A higher percentile means a stronger position within that category's ranked cohort; 100 is highest.
- Inst. Following13%Eligible cohort rank #27 of 31Category score 39.5
- Agentic48%Eligible cohort rank #62 of 119Category score 47.5
Category evidence
Scores and ranks appear only where this model has published benchmark evidence. Categories without displayable source records remain not measured.
| Category | Score | Rank | Percentile | Weight | Benchmarks | Evidence |
|---|---|---|---|---|---|---|
| AgenticRank #62 of 119Percentile 48thWeight 22%1 benchmarkVerified | 47.5 | #62 of 119 | 48th | 22% | 1 benchmark | Verified |
| CodingWeight 20%0 benchmarksNot measured | Not measured | Not ranked | Not available | 20% | 0 benchmarks | Not measured |
| ReasoningWeight 17%0 benchmarksNot measured | Not measured | Not ranked | Not available | 17% | 0 benchmarks | Not measured |
| KnowledgeRank Not rankedWeight 12%3 benchmarksVerified | 46.4 | Not ranked | Not available | 12% | 3 benchmarks | Verified |
| MathWeight 5%0 benchmarksNot measured | Not measured | Not ranked | Not available | 5% | 0 benchmarks | Not measured |
| MultilingualWeight 7%0 benchmarksNot measured | Not measured | Not ranked | Not available | 7% | 0 benchmarks | Not measured |
| MultimodalWeight 12%0 benchmarksNot measured | Not measured | Not ranked | Not available | 12% | 0 benchmarks | Not measured |
| Inst. FollowingRank #27 of 31Percentile 13thWeight 5%1 benchmarkVerified | 39.5 | #27 of 31 | 13th | 5% | 1 benchmark | Verified |
Benchmark Details
Rows below have a displayable published verification record. Each source link and provenance note remains in the page HTML while its category is closed. Source-unverified manual rows and generated rows stay hidden.
Agentic1 benchmark
Berkeley Function Calling Leaderboard v4
Knowledge3 benchmarks
Graduate-Level Google-Proof Q&A
GPQA Diamond
Inst. Following1 benchmark
Instruction-Following Eval
Frequently Asked Questions
How does Mellum2-12B-A2.5B-Thinking perform overall in AI benchmarks?
Mellum2-12B-A2.5B-Thinking has 5 published benchmark scores on BenchLM, but it does not yet have enough non-generated coverage to receive a global overall rank.
Is Mellum2-12B-A2.5B-Thinking good for knowledge and understanding?
Mellum2-12B-A2.5B-Thinking has visible benchmark coverage in knowledge and understanding, but BenchLM does not currently assign it a global category rank there.
Is Mellum2-12B-A2.5B-Thinking good for agentic tool use and computer tasks?
Mellum2-12B-A2.5B-Thinking ranks #62 out of 119 models in agentic tool use and computer tasks benchmarks with an average score of 47.5. There are stronger options in this category.
Is Mellum2-12B-A2.5B-Thinking good for instruction following?
Mellum2-12B-A2.5B-Thinking ranks #27 out of 31 models in instruction following benchmarks with an average score of 39.5. There are stronger options in this category.
Is Mellum2-12B-A2.5B-Thinking open source?
Yes, Mellum2-12B-A2.5B-Thinking is an open weight model created by JetBrains, meaning it can be downloaded and run locally or fine-tuned for specific use cases.
Which sibling models are related to Mellum2-12B-A2.5B-Thinking?
Mellum2-12B-A2.5B-Thinking belongs to the Mellum2 12B-A2.5B family. Related variants on BenchLM include Mellum2-12B-A2.5B-Instruct.
Does Mellum2-12B-A2.5B-Thinking have full benchmark coverage on BenchLM?
Not yet. Mellum2-12B-A2.5B-Thinking currently has 5 published benchmark scores out of the 323 benchmarks BenchLM tracks. BenchLM only exposes non-generated public benchmark rows, so missing categories stay blank until a sourced evaluation is available.
What is the context window size of Mellum2-12B-A2.5B-Thinking?
Mellum2-12B-A2.5B-Thinking has a published context window of 128K, which determines how much text it can process in a single interaction.
Related Resources
Choose with this week’s evidence
Join 2,000+ readers for ranking moves, new releases, pricing changes, and the evidence behind them.
Free. One email per week.