Model profile
Trinity-Large-Thinking
Evidence coverage
19 of 323 tracked benchmarks are published. 1 is verified and 18 provisional. 7 of 8 categories are measured.
- Published / tracked
- 19 / 323
- Verified
- 1
- Provisional
- 18
- Categories with evidence
- 7 / 8
Evidence by category
- Agentic4 benchmarksMixed evidence
- Coding2 benchmarksReported
- Reasoning2 benchmarksReported
- Knowledge8 benchmarksReported
- Math1 benchmarkReported
- Multilingual0 benchmarksNot measured
- Multimodal1 benchmarkReported
- Inst. Following1 benchmarkReported
Trinity-Large-Thinking ranks #100 out of 200 models on the public leaderboard with an overall score of 52.32/100. It also ranks #55 out of 99 on the verified leaderboard. While not a frontier model, it offers specific advantages depending on the use case.
Trinity-Large-Thinking is a open weight model with a 512K token context window. It uses explicit chain-of-thought reasoning, which typically improves performance on math and complex reasoning tasks at the cost of higher latency and token usage.
Trinity-Large-Thinking sits inside the Trinity Large family alongside Trinity-Large-Preview. This profile currently has 19 of 323 tracked benchmarks. BenchLM only exposes non-generated benchmark rows publicly, so missing categories stay blank until a sourced evaluation is available.
Its strongest category is Coding (#87). This performance profile makes it particularly well-suited for software development and code generation tasks.
Peer position
Exact provisional scores and ranks for the closest listed peers. A score can appear before a model clears the evidence threshold for a rank, so equal scores can have different rank states.
Range 51.47–52.94
- Nemotron 3 Nano 30BNVIDIACompare#9852.94Nemotron 3 Nano 30B is #98 with a score of 52.94.
- GPT-5.1-CodexOpenAICompare#9952.72GPT-5.1-Codex is #99 with a score of 52.72.
- Trinity-Large-ThinkingCurrent modelArcee AI#10052.32Trinity-Large-Thinking is #100 with a score of 52.32.
- Qwen2.5-72BAlibabaCompare#10152.15Qwen2.5-72B is #101 with a score of 52.15.
- Llama 3.1 405BMetaCompare#10251.71Llama 3.1 405B is #102 with a score of 51.71.
- DeepSeek-R1DeepSeekCompare#10351.67DeepSeek-R1 is #103 with a score of 51.67.
- Qwen3.6-35B-A3BAlibabaCompare#10451.47Qwen3.6-35B-A3B is #104 with a score of 51.47.
Category percentile
More
Relative position among models eligible for each sourced category. A higher percentile means a stronger position within that category's ranked cohort; 100 is highest.
- Coding29%Eligible cohort rank #87 of 122Category score 46.6
Category evidence
Scores and ranks appear only where this model has published benchmark evidence. Categories without displayable source records remain not measured.
| Category | Score | Rank | Percentile | Weight | Benchmarks | Evidence |
|---|---|---|---|---|---|---|
| AgenticWeight 22%4 benchmarksMixed sources | Score pending | Not ranked | Not available | 22% | 4 benchmarks | Mixed sources |
| CodingRank #87 of 122Percentile 29thWeight 20%2 benchmarksReported | 46.6 | #87 of 122 | 29th | 20% | 2 benchmarks | Reported |
| ReasoningWeight 17%2 benchmarksReported | Score pending | Not ranked | Not available | 17% | 2 benchmarks | Reported |
| KnowledgeWeight 12%8 benchmarksReported | Score pending | Not ranked | Not available | 12% | 8 benchmarks | Reported |
| MathWeight 5%1 benchmarkReported | Score pending | Not ranked | Not available | 5% | 1 benchmark | Reported |
| MultilingualWeight 7%0 benchmarksNot measured | Not measured | Not ranked | Not available | 7% | 0 benchmarks | Not measured |
| MultimodalWeight 12%1 benchmarkReported | Score pending | Not ranked | Not available | 12% | 1 benchmark | Reported |
| Inst. FollowingRank Not rankedWeight 5%1 benchmarkReported | 52.3 | Not ranked | Not available | 5% | 1 benchmark | Reported |
Chatbot Arena performance
Scroll horizontally to inspect confidence intervals and vote counts.
| View | Elo | Confidence interval | Votes |
|---|---|---|---|
| Text Overall | 1369 | ±4.7 | 29,139 |
| Coding | 1414 | ±7.6 | 8,159 |
| Math | 1384 | ±15.2 | 1,618 |
| Instruction Following | 1358 | ±7.0 | 9,786 |
| Creative Writing | 1336 | ±9.6 | 4,548 |
| Multi-turn | 1349 | ±9.1 | 5,385 |
| Hard Prompts | 1386 | ±5.7 | 19,102 |
| Hard Prompts (English) | 1396 | ±7.1 | 9,515 |
| Longer Query | 1367 | ±6.9 | 12,281 |
Benchmark Details
Rows below have a displayable published verification record. Each source link and provenance note remains in the page HTML while its category is closed. Source-unverified manual rows and generated rows stay hidden.
Agentic4 benchmarks
τ²-Bench Tool-Agent-User Evaluation
Gert Labs Composite Game Benchmark
GDPval-AA normalized
Coding2 benchmarks
SWE-bench Verified (mini-swe-agent-v2)
Artificial Analysis SciCode
Reasoning2 benchmarks
Artificial Analysis Long Context Reasoning
Critical Physics Tasks
Knowledge8 benchmarks
GPQA Diamond
MMLU-Pro first-party comparison snapshot
Artificial Analysis GPQA Diamond
Artificial Analysis Humanity's Last Exam
Artificial Analysis Omniscience Index
Artificial Analysis Omniscience Accuracy
Artificial Analysis Omniscience Hallucination Rate
Math1 benchmark
AIME25 first-party comparison snapshot
Multimodal1 benchmark
Design Arena Website Elo
Inst. Following1 benchmark
Artificial Analysis IFBench
Frequently Asked Questions
How does Trinity-Large-Thinking perform overall in AI benchmarks?
Trinity-Large-Thinking has 19 published benchmark scores on BenchLM, but it does not yet have enough non-generated coverage to receive a global overall rank.
Is Trinity-Large-Thinking good for knowledge and understanding?
Trinity-Large-Thinking has visible benchmark coverage in knowledge and understanding, but BenchLM does not currently assign it a global category rank there.
Is Trinity-Large-Thinking good for coding and programming?
Trinity-Large-Thinking ranks #87 out of 122 models in coding and programming benchmarks with an average score of 46.6. There are stronger options in this category.
Is Trinity-Large-Thinking good for mathematics?
Trinity-Large-Thinking has visible benchmark coverage in mathematics, but BenchLM does not currently assign it a global category rank there.
Is Trinity-Large-Thinking good for reasoning and logic?
Trinity-Large-Thinking has visible benchmark coverage in reasoning and logic, but BenchLM does not currently assign it a global category rank there.
Is Trinity-Large-Thinking good for agentic tool use and computer tasks?
Trinity-Large-Thinking has visible benchmark coverage in agentic tool use and computer tasks, but BenchLM does not currently assign it a global category rank there.
Is Trinity-Large-Thinking good for multimodal and grounded tasks?
Trinity-Large-Thinking has visible benchmark coverage in multimodal and grounded tasks, but BenchLM does not currently assign it a global category rank there.
Is Trinity-Large-Thinking good for instruction following?
Trinity-Large-Thinking has visible benchmark coverage in instruction following, but BenchLM does not currently assign it a global category rank there.
Is Trinity-Large-Thinking open source?
Yes, Trinity-Large-Thinking is an open weight model created by Arcee AI, meaning it can be downloaded and run locally or fine-tuned for specific use cases.
Which sibling models are related to Trinity-Large-Thinking?
Trinity-Large-Thinking belongs to the Trinity Large family. Related variants on BenchLM include Trinity-Large-Preview.
Does Trinity-Large-Thinking have full benchmark coverage on BenchLM?
Not yet. Trinity-Large-Thinking currently has 19 published benchmark scores out of the 323 benchmarks BenchLM tracks. BenchLM only exposes non-generated public benchmark rows, so missing categories stay blank until a sourced evaluation is available.
What is the context window size of Trinity-Large-Thinking?
Trinity-Large-Thinking has a published context window of 512K, which determines how much text it can process in a single interaction.
Related Resources
Choose with this week’s evidence
Join 2,000+ readers for ranking moves, new releases, pricing changes, and the evidence behind them.
Free. One email per week.