Skip to main content

BenchLM recommendation

Best LLMs for Writing in 2026

Data verified

As of July 23, 2026, the top model in best llms for writing on the BenchLM leaderboard is MAI-Thinking-1 with a score of 97.8.

Last verified: July 23, 2026

There is no single "writing benchmark," so BenchLM ranks writing capability by the instruction-following category — the best available proxy for whether a model matches your tone, structure, and length constraints — read alongside Arena Elo, the human-preference signal. Claude models hold the top Arena Elo scores; the instruction-following table below shows who follows a brief most reliably.

Unless noted otherwise, ranking surfaces on this page use BenchLM's provisional leaderboard lane rather than the stricter sourced-only verified leaderboard.

Bottom line: instruction following is what separates writing models in practice. Claude Fable 5 holds the top Arena Elo (1508) for human preference; the instruction-following leaders below execute a brief most faithfully.

MAI-Thinking-1 leads this ranking with a score of 97.8, followed by Qwen3.5-27B (93) and Agents-A1 (93). There is meaningful separation between the top models, suggesting genuine performance differences.

The best open-weight option is Qwen3.5-27B (ranked #2 with a score of 93). Open-weight models are highly competitive in this category — self-hosting is a viable alternative to proprietary APIs.

This ranking is based on provisional weighted averages across the scoring benchmarks in instructionFollowing tracked by BenchLM.ai. For detailed model profiles, click any model name below. To compare two specific models head-to-head, use the "vs #" links.

What changed

MAI-Thinking-1 leads the instruction-following category at 93.3.

GPT-5.4 close second at 92.6 — the reliable structured-writing workhorse.

Claude Opus 4.6 strong instruction following (91.3) with Claude's prose quality.

How to choose

Full Rankings (31 models)

1
MAI-Thinking-1
Microsoft·Proprietary·256K

97.8

prov. avg

2
Qwen3.5-27B
Alibaba·Open Weight·262K

93

prov. avg

3
Agents-A1
InternScience·Open Weight·262K

93

prov. avg

4
Nemotron 3 Ultra
NVIDIA·Open Weight·1M

91.9

prov. avg

5
Qwen3.7 Plus
Alibaba·Proprietary·1M

91.4

prov. avg

6
Grok 4.3
xAI·Proprietary·1M

91.2

prov. avg

7
Qwen3.7 Max
Alibaba·Proprietary·1M

91

prov. avg

8
Kimi K2.5
Moonshot AI·Open Weight·256K

90.3

prov. avg

9
o3-mini
OpenAI·Proprietary·200K

90.3

prov. avg

10
Qwen3.5-122B-A10B
Alibaba·Open Weight·262K

88.9

prov. avg

11
Inkling
Thinking Machines Lab·Open Weight·1M

88.6

prov. avg

12
GLM-5
Z.AI·Open Weight·200K

86.5

prov. avg

13
Qwen3.5 397B
Alibaba·Open Weight·128K

86.5

prov. avg

14
Qwen3.6 Plus
Alibaba·Proprietary·1M

86.4

prov. avg

15
o1
OpenAI·Proprietary·200K

85.4

prov. avg

16
Qwen3.5-35B-A3B
Alibaba·Open Weight·262K

84.5

prov. avg

17
Gemini 3.5 Flash
Google·Proprietary·1M

82.4

prov. avg

18
Nemotron 3 Nano Omni 30B A3B
NVIDIA·Open Weight·256K

78.7

prov. avg

19
GPT-4.1 mini
OpenAI·Proprietary·1M

74.6

prov. avg

20
GPT-4.1
OpenAI·Proprietary·1M

71.4

prov. avg

21
GPT-4.1 nano
OpenAI·Proprietary·1M

59.1

prov. avg

22
Hy3 Preview
Tencent·Open Weight·256K

59

prov. avg

23
Claude Opus 4.5
Anthropic·Proprietary·200K

56.4

prov. avg

24
LFM2.5-8B-A1B
LiquidAI·Open Weight·128K

55.7

prov. avg

25
Ling 2.6 Flash
InclusionAI·Open Weight·262K

48.2

prov. avg

26
ZAYA1-8B
Zyphra·Open Weight·131K

40.8

prov. avg

27
Mellum2-12B-A2.5B-Thinking
JetBrains·Open Weight·128K

39.5

prov. avg

28
Mellum2-12B-A2.5B-Instruct
JetBrains·Open Weight·128K

37.5

prov. avg

29
LFM2.5-VL-450M
LiquidAI·Open Weight·128K

25.5

prov. avg

30
MiniCPM5-1B
OpenBMB·Open Weight·131K

24.7

prov. avg

31
LFM2.5-230M
LiquidAI·Open Weight·32K

1

prov. avg

Key Takeaways

The top model is MAI-Thinking-1 by Microsoft with a provisional score of 97.8.

The best open-weight model is Qwen3.5-27B at position #2.

31 models are included in this ranking.

Score in Context

What these scores mean

Writing quality has no direct benchmark, so this page ranks by instruction following — whether the model does what the brief asked. Read it with Arena Elo, the blind human-preference score, for prose quality.

Known limitations

Style is subjective and prompt-sensitive. Instruction-following scores reward constraint compliance, not voice — a model can follow your brief perfectly and still write flat prose. Test your actual editing workflow.

Best LLMs for Writing FAQ

What is the best LLM for writing?

Claude Fable 5 is the strongest overall pick — it holds the highest Arena Elo BenchLM tracks (1508), the best human-preference signal for prose, with top-tier instruction following. For structured copy at lower cost, GPT-5.4's 92.6 instruction-following score makes it the workhorse choice.

What is the best LLM for creative writing?

Human preference is the best available signal for creative prose, and Claude models lead it: Claude Fable 5 holds the top Arena Elo on BenchLM's board. Non-reasoning models often feel more natural for iterative drafting because they respond without a thinking pause.

What is the best AI for resume writing?

Resume writing is an instruction-following task — strict format, tight length, specific tone. Any model above ~90 in the table handles it well; the practical differences are price and speed, so a mid-tier model like GPT-5.4 or Claude Sonnet 5 is the sensible default.

Are benchmarks meaningful for writing quality?

Partially. Instruction following measures whether the model obeyed the brief — essential for professional writing — and Arena Elo captures blind human preference. Neither measures your voice. Use the scores to shortlist, then run a 10-prompt bake-off in your own editing workflow.

Last updated: July 23, 2026

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.