Model profile
Qwen3.5 Flash
Evidence coverage
2 of 323 tracked benchmarks are published. 2 are verified and 0 provisional. 1 of 8 categories are measured.
- Published / tracked
- 2 / 323
- Verified
- 2
- Provisional
- 0
- Categories with evidence
- 1 / 8
Evidence by category
- Agentic0 benchmarksNot measured
- Coding0 benchmarksNot measured
- Reasoning0 benchmarksNot measured
- Knowledge0 benchmarksNot measured
- Math2 benchmarksVerified
- Multilingual0 benchmarksNot measured
- Multimodal0 benchmarksNot measured
- Inst. Following0 benchmarksNot measured
Qwen3.5 Flash ranks #134 out of 200 models on the public leaderboard with an overall score of 47.65/100. It also ranks #69 out of 99 on the verified leaderboard. While not a frontier model, it offers specific advantages depending on the use case.
Qwen3.5 Flash is a proprietary model with a 1M token context window. It uses explicit chain-of-thought reasoning, which typically improves performance on math and complex reasoning tasks at the cost of higher latency and token usage.
This profile currently has 2 of 323 tracked benchmarks. BenchLM only exposes non-generated benchmark rows publicly, so missing categories stay blank until a sourced evaluation is available.
Peer position
Exact provisional scores and ranks for the closest listed peers. A score can appear before a model clears the evidence threshold for a rank, so equal scores can have different rank states.
Range 47.29–47.89
- o3OpenAICompare#13147.89o3 is #131 with a score of 47.89.
- Claude 3.5 SonnetAnthropicCompare#13247.74Claude 3.5 Sonnet is #132 with a score of 47.74.
- GLM-4.5-AirZ.AICompare#13347.7GLM-4.5-Air is #133 with a score of 47.7.
- Qwen3.5 FlashCurrent modelAlibaba#13447.65Qwen3.5 Flash is #134 with a score of 47.65.
- Command A+CohereCompare#13547.51Command A+ is #135 with a score of 47.51.
- o3-miniOpenAICompare#13647.41o3-mini is #136 with a score of 47.41.
- Gemma 4 12BGoogleCompare#13747.29Gemma 4 12B is #137 with a score of 47.29.
Category percentile
More
Relative position among models eligible for each sourced category. A higher percentile means a stronger position within that category's ranked cohort; 100 is highest.
Category evidence
Scores and ranks appear only where this model has published benchmark evidence. Categories without displayable source records remain not measured.
| Category | Score | Rank | Percentile | Weight | Benchmarks | Evidence |
|---|---|---|---|---|---|---|
| AgenticWeight 22%0 benchmarksNot measured | Not measured | Not ranked | Not available | 22% | 0 benchmarks | Not measured |
| CodingWeight 20%0 benchmarksNot measured | Not measured | Not ranked | Not available | 20% | 0 benchmarks | Not measured |
| ReasoningWeight 17%0 benchmarksNot measured | Not measured | Not ranked | Not available | 17% | 0 benchmarks | Not measured |
| KnowledgeWeight 12%0 benchmarksNot measured | Not measured | Not ranked | Not available | 12% | 0 benchmarks | Not measured |
| MathRank Not rankedWeight 5%2 benchmarksVerified | 28.7 | Not ranked | Not available | 5% | 2 benchmarks | Verified |
| MultilingualWeight 7%0 benchmarksNot measured | Not measured | Not ranked | Not available | 7% | 0 benchmarks | Not measured |
| MultimodalWeight 12%0 benchmarksNot measured | Not measured | Not ranked | Not available | 12% | 0 benchmarks | Not measured |
| Inst. FollowingWeight 5%0 benchmarksNot measured | Not measured | Not ranked | Not available | 5% | 0 benchmarks | Not measured |
Chatbot Arena performance
Scroll horizontally to inspect confidence intervals and vote counts.
| View | Elo | Confidence interval | Votes |
|---|---|---|---|
| Text Overall | 1397 | ±3.9 | 55,880 |
| Coding | 1438 | ±6.0 | 15,930 |
| Math | 1403 | ±11.2 | 3,067 |
| Instruction Following | 1383 | ±5.6 | 17,844 |
| Creative Writing | 1339 | ±7.4 | 8,627 |
| Multi-turn | 1396 | ±7.0 | 9,820 |
| Hard Prompts | 1416 | ±4.6 | 35,854 |
| Hard Prompts (English) | 1422 | ±5.8 | 17,110 |
| Longer Query | 1402 | ±5.5 | 22,324 |
Benchmark Details
Rows below have a displayable published verification record. Each source link and provenance note remains in the page HTML while its category is closed. Source-unverified manual rows and generated rows stay hidden.
Math2 benchmarks
FrontierMath v2 Tiers 1-3
FrontierMath v2 Tier 4
Frequently Asked Questions
How does Qwen3.5 Flash perform overall in AI benchmarks?
Qwen3.5 Flash has 2 published benchmark scores on BenchLM, but it does not yet have enough non-generated coverage to receive a global overall rank.
Is Qwen3.5 Flash good for mathematics?
Qwen3.5 Flash has visible benchmark coverage in mathematics, but BenchLM does not currently assign it a global category rank there.
Does Qwen3.5 Flash have full benchmark coverage on BenchLM?
Not yet. Qwen3.5 Flash currently has 2 published benchmark scores out of the 323 benchmarks BenchLM tracks. BenchLM only exposes non-generated public benchmark rows, so missing categories stay blank until a sourced evaluation is available.
What is the context window size of Qwen3.5 Flash?
Qwen3.5 Flash has a published context window of 1M, which determines how much text it can process in a single interaction.
Related Resources
Choose with this week’s evidence
Join 2,000+ readers for ranking moves, new releases, pricing changes, and the evidence behind them.
Free. One email per week.