Skip to main content

Model profile

Qwen 3.6 Max (preview)

AlibabaCurrentReleased Apr 20, 2026
Data verified
Overall Score
59.72Public #53 of 200Verified #40 of 99
Arena Elo
1460
Eligible category ranks
2of 8
Price (1M tokens)
Not listedAPI pricing
Speed
Not listed
Context
256K

Evidence coverage

21 of 323 tracked benchmarks are published. 10 are verified and 11 provisional. 6 of 8 categories are measured.

Updated July 23, 2026Methodology
Published / tracked
21 / 323
Verified
10
Provisional
11
Categories with evidence
6 / 8

Evidence by category

  • Agentic4 benchmarks
    Mixed evidence
  • Coding5 benchmarks
    Mixed evidence
  • Reasoning2 benchmarks
    Reported
  • Knowledge7 benchmarks
    Mixed evidence
  • Math2 benchmarks
    Verified
  • Multilingual0 benchmarks
    Not measured
  • Multimodal0 benchmarks
    Not measured
  • Inst. Following1 benchmark
    Reported
ProprietaryReasoning
Confidence:
Medium
preview

Qwen 3.6 Max (preview) ranks #53 out of 200 models on the public leaderboard with an overall score of 59.72/100. It also ranks #40 out of 99 on the verified leaderboard. While not a frontier model, it offers specific advantages depending on the use case.

Qwen 3.6 Max (preview) is a proprietary model with a 256K token context window. It uses explicit chain-of-thought reasoning, which typically improves performance on math and complex reasoning tasks at the cost of higher latency and token usage.

This profile currently has 21 of 323 tracked benchmarks. BenchLM only exposes non-generated benchmark rows publicly, so missing categories stay blank until a sourced evaluation is available.

Its strongest category is Agentic (#18), while its weakest is Coding (#26). This performance profile makes it particularly useful for coding agents, browser research, and computer-use workflows.

Peer position

Exact provisional scores and ranks for the closest listed peers. A score can appear before a model clears the evidence threshold for a rank, so equal scores can have different rank states.

Range 59.3559.97

  1. Grok 4.1
    xAI
    #5159.97
    Grok 4.1 is #51 with a score of 59.97.
    Compare
  2. GLM-5 (Reasoning)
    Z.AI
    #5259.77
    GLM-5 (Reasoning) is #52 with a score of 59.77.
    Compare
  3. Qwen 3.6 Max (preview)Current model
    Alibaba
    #5359.72
    Qwen 3.6 Max (preview) is #53 with a score of 59.72.
  4. Kimi K2.5
    Moonshot AI
    #5459.66
    Kimi K2.5 is #54 with a score of 59.66.
    Compare
  5. MiniMax M2.5
    MiniMax
    #5559.52
    MiniMax M2.5 is #55 with a score of 59.52.
    Compare
  6. Qwen3.5 397B (Reasoning)
    Alibaba
    #5659.5
    Qwen3.5 397B (Reasoning) is #56 with a score of 59.5.
    Compare
  7. Kimi K2.5 (Reasoning)
    Moonshot AI
    #5759.35
    Kimi K2.5 (Reasoning) is #57 with a score of 59.35.
    Compare

Category percentile

More

Relative position among models eligible for each sourced category. A higher percentile means a stronger position within that category's ranked cohort; 100 is highest.

  1. Agentic86%
    Eligible cohort rank #18 of 119Category score 57.2
  2. Coding79%
    Eligible cohort rank #26 of 122Category score 58.9

Category evidence

Scores and ranks appear only where this model has published benchmark evidence. Categories without displayable source records remain not measured.

Category scores, ranks, weighting, benchmark coverage, and evidence status
CategoryScore
AgenticRank #18 of 119Percentile 86thWeight 22%4 benchmarksMixed sources57.2
CodingRank #26 of 122Percentile 79thWeight 20%5 benchmarksMixed sources58.9
ReasoningWeight 17%2 benchmarksReportedScore pending
KnowledgeRank Not rankedWeight 12%7 benchmarksMixed sources68.4
MathRank Not rankedWeight 5%2 benchmarksVerified41.8
MultilingualWeight 7%0 benchmarksNot measuredNot measured
MultimodalWeight 12%0 benchmarksNot measuredNot measured
Inst. FollowingWeight 5%1 benchmarkReportedScore pending

Chatbot Arena performance

Scroll horizontally to inspect confidence intervals and vote counts.

Chatbot Arena Elo, confidence interval, and vote count by evaluation view
ViewEloConfidence intervalVotes
Text Overall1460±8.45,191
Coding1508±15.61,547
Math1477±30.0359
Instruction Following1450±14.21,761
Creative Writing1436±23.1716
Multi-turn1469±20.4895
Hard Prompts1481±10.43,392
Hard Prompts (English)1483±14.91,637
Longer Query1475±13.62,126

Benchmark Details

Rows below have a displayable published verification record. Each source link and provenance note remains in the page HTML while its category is closed. Source-unverified manual rows and generated rows stay hidden.

Agentic4 benchmarks
Terminal-Bench 2.0Provider exact
65.4%Weighted 38%
Source: Qwen: Qwen3.6-Max-Preview launch pageProvenance: Provider exact
QwenClawBenchProvider exact
59.0%Display only
Source: Qwen: Qwen3.6-Max-Preview launch pageProvenance: First-party launch chart shared on the Qwen3.6-Max-Preview page reports QwenClawBench at 59.0.
QwenWebBenchProvider exact
1532Display only
Source: Qwen: Qwen3.6-Max-Preview launch pageProvenance: First-party launch chart shared on the Qwen3.6-Max-Preview page reports QwenWebBench at 1532 Elo.
τ²-bench resultsReported

τ²-Bench Tool-Agent-User Evaluation

95.9%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Coding5 benchmarks
SciCodeProvider exact

Scientific Code Benchmark

47%Weighted 16%
Source: Qwen: Qwen3.6-Max-Preview launch pageProvenance: First-party launch chart shared on the Qwen3.6-Max-Preview page reports SciCode at 47.0.
SWE-bench ProProvider exact
57.3%Weighted 10%
Source: Qwen: Qwen3.6-Max-Preview launch pageProvenance: First-party launch chart shared on the Qwen3.6-Max-Preview page reports SWE-bench Pro at 57.3.
NL2RepoProvider exact
42.9%Display only
Source: Qwen: Qwen3.6-Max-Preview launch pageProvenance: First-party launch chart shared on the Qwen3.6-Max-Preview page reports NL2Repo at 42.9.
Terminal-Bench 2.0Provider exact
65.4%Display only
Source: Qwen: Qwen3.6-Max-Preview launch pageProvenance: First-party launch chart shared on the Qwen3.6-Max-Preview page reports Terminal-Bench 2.0 (Terminus-2) at 65.4.
AA-SciCodeReported

Artificial Analysis SciCode

46.9%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Reasoning2 benchmarks
AA-LCRReported

Artificial Analysis Long Context Reasoning

69.7%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
CritPtReported

Critical Physics Tasks

3.7%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Knowledge7 benchmarks
SuperGPQAProvider exact

SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines

73.9%Weighted 7%
Source: Qwen: Qwen3.6-Max-Preview launch pageProvenance: First-party launch chart shared on the Qwen3.6-Max-Preview page reports SuperGPQA at 73.9.
Artificial Analysis Intelligence IndexReported
40.0%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-GPQA DiamondReported

Artificial Analysis GPQA Diamond

88.8%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-HLEReported

Artificial Analysis Humanity's Last Exam

28.9%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-Omniscience IndexReported

Artificial Analysis Omniscience Index

10.2%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-Omniscience AccuracyReported

Artificial Analysis Omniscience Accuracy

37.7%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-Omniscience Hallucination RateReported

Artificial Analysis Omniscience Hallucination Rate

44.2%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Math2 benchmarks
FrontierMath v2 (Tiers 1-3)Benchmark exact

FrontierMath v2 Tiers 1-3

23.103%Weighted 30%
Source: Epoch AI FrontierMath v2 leaderboardProvenance: Epoch AI reports FrontierMath v2 Tiers 1-3 at 23.103% for qwen3.6-max-preview. BenchLM selects the highest published thinking effort for the model and stores the v2 benchmark slice separately.
FrontierMath v2 (Tier 4)Benchmark exact

FrontierMath v2 Tier 4

4.167%Weighted 10%
Source: Epoch AI FrontierMath v2 leaderboardProvenance: Epoch AI reports FrontierMath v2 Tier 4 at 4.167% for qwen3.6-max-preview. BenchLM selects the highest published thinking effort for the model and stores the v2 benchmark slice separately.
Inst. Following1 benchmark
AA-IFBenchReported

Artificial Analysis IFBench

76.6%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.

Frequently Asked Questions

How does Qwen 3.6 Max (preview) perform overall in AI benchmarks?

Qwen 3.6 Max (preview) has 21 published benchmark scores on BenchLM, but it does not yet have enough non-generated coverage to receive a global overall rank.

Is Qwen 3.6 Max (preview) good for knowledge and understanding?

Qwen 3.6 Max (preview) has visible benchmark coverage in knowledge and understanding, but BenchLM does not currently assign it a global category rank there.

Is Qwen 3.6 Max (preview) good for coding and programming?

Qwen 3.6 Max (preview) ranks #26 out of 122 models in coding and programming benchmarks with an average score of 58.9. There are stronger options in this category.

Is Qwen 3.6 Max (preview) good for mathematics?

Qwen 3.6 Max (preview) has visible benchmark coverage in mathematics, but BenchLM does not currently assign it a global category rank there.

Is Qwen 3.6 Max (preview) good for reasoning and logic?

Qwen 3.6 Max (preview) has visible benchmark coverage in reasoning and logic, but BenchLM does not currently assign it a global category rank there.

Is Qwen 3.6 Max (preview) good for agentic tool use and computer tasks?

Qwen 3.6 Max (preview) ranks #18 out of 119 models in agentic tool use and computer tasks benchmarks with an average score of 57.2. There are stronger options in this category.

Is Qwen 3.6 Max (preview) good for instruction following?

Qwen 3.6 Max (preview) has visible benchmark coverage in instruction following, but BenchLM does not currently assign it a global category rank there.

Does Qwen 3.6 Max (preview) have full benchmark coverage on BenchLM?

Not yet. Qwen 3.6 Max (preview) currently has 21 published benchmark scores out of the 323 benchmarks BenchLM tracks. BenchLM only exposes non-generated public benchmark rows, so missing categories stay blank until a sourced evaluation is available.

What is the context window size of Qwen 3.6 Max (preview)?

Qwen 3.6 Max (preview) has a published context window of 256K, which determines how much text it can process in a single interaction.

Last updated: July 23, 2026 · Runtime metrics stay blank until BenchLM has a sourced snapshot.

Choose with this week’s evidence

Join 2,000+ readers for ranking moves, new releases, pricing changes, and the evidence behind them.

Free. One email per week.