Skip to main content

LLM Speed & Latency Comparison

Compare inference speed across every major AI model. Tokens/sec measures output generation speed. Latency measures time to the first answer token — for reasoning models this includes thinking time, so it reflects end-to-end response latency rather than raw TTFT.

Speed data from Artificial Analysis. Last updated: 2026-07-21. Median tokens/s, Latency first answer chunk (s).

Fastest Output

Mercury 2

789 tok/s · Inception

Lowest Latency

Command A+

0.25s to first answer · Cohere

Fastest (Score 70+)

GPT-5.4

74 tok/s · Score: 74.24

Top 15 — Output Speed (tok/s)

Ultra FastFastMediumSlow
  1. Mercury 2789 tok/s, Inception, 3.88 seconds first-answer latency, score 51.28
  2. Nemotron 3 Super 100B367 tok/s, NVIDIA, 0.71 seconds first-answer latency, score 50.08
  3. GPT-OSS 20B313 tok/s, OpenAI, 0.65 seconds first-answer latency, score 42.74
  4. Gemini 3.5 Flash284.2 tok/s, Google, 18.55 seconds first-answer latency, score 64.75
  5. Ministral 3 3B274 tok/s, Mistral, 0.42 seconds first-answer latency, score 18.27
  6. Command A+272 tok/s, Cohere, 0.25 seconds first-answer latency, score 47.51
  7. GPT-OSS 120B262 tok/s, OpenAI, 0.79 seconds first-answer latency, score 50.08
  8. Grok 4.20233 tok/s, xAI, 10.33 seconds first-answer latency, score 54.68
  9. Gemini 2.5 Flash221 tok/s, Google, 0.5 seconds first-answer latency, score 48.09
  10. Ling 2.6 Flash209.5 tok/s, InclusionAI, 1.07 seconds first-answer latency, score 43.87
  11. Grok 4.3209 tok/s, xAI, 12.36 seconds first-answer latency, score 65.1
  12. Gemini 3.1 Flash-Lite205 tok/s, Google, 7.5 seconds first-answer latency, score 50.83
  13. GPT-5.4 mini201 tok/s, OpenAI, 3.85 seconds first-answer latency, score 56.77
  14. GPT-5.4 nano191 tok/s, OpenAI, 3.64 seconds first-answer latency, score 66.79
  15. Grok 3 Mini190 tok/s, xAI, 0.54 seconds first-answer latency

Bar length is relative to the fastest model in the current filtered view. Exact median speed is shown at right; open a model for its complete evidence profile.

Average Speed by Provider

NVIDIA

260 tok/s avg · 2 models

1.3s avg latency

xAI

166 tok/s avg · 6 models

7s avg latency

Google

154 tok/s avg · 8 models

14.2s avg latency

Mistral

126 tok/s avg · 7 models

0.8s avg latency

OpenAI

121 tok/s avg · 25 models

42.3s avg latency

Meta

93 tok/s avg · 3 models

1.3s avg latency

Z.AI

82 tok/s avg · 5 models

1.3s avg latency

Anthropic

52 tok/s avg · 7 models

3.3s avg latency

DeepSeek

48 tok/s avg · 2 models

2.3s avg latency

MiniMax

46 tok/s avg · 2 models

2.3s avg latency

Moonshot AI

44 tok/s avg · 2 models

1.9s avg latency

78 models match these filters.

RankModelSpeed / latency
1
Mercury 2Inception · score 51.28
789 tok/s3.88s latency
3
GPT-OSS 20BOpenAI · score 42.74
313 tok/s0.65s latency
6
Command A+Cohere · score 47.51
272 tok/s0.25s latency
8
Grok 4.20xAI · score 54.68
233 tok/s10.33s latency
11
Grok 4.3xAI · score 65.1
209 tok/s12.36s latency
15
Grok 3 MinixAI · score not scored
190 tok/s0.54s latency
21
o3-miniOpenAI · score 47.41
160 tok/s7.12s latency
24
Nova ProAmazon · score 20.32
141 tok/s0.81s latency
27
GPT-5 nanoOpenAI · score 46.36
137 tok/s83.3s latency
28
GPT-4oOpenAI · score 41.49
131 tok/s0.81s latency
33
o3OpenAI · score 47.89
118 tok/s5.38s latency
35
GPT-5.1OpenAI · score 53.65
111 tok/s57.47s latency
39
GPT-4.1OpenAI · score 51.11
108 tok/s1.02s latency
41
o1OpenAI · score 48.1
98 tok/s32.29s latency
46
GPT-5 miniOpenAI · score 43.93
86 tok/s65.32s latency
49
GLM-4.7Z.AI · score 61.16
82 tok/s1.1s latency
52
GPT-5.4OpenAI · score 74.24
74 tok/s151.79s latency
53
GLM-5Z.AI · score 66.06
74 tok/s1.64s latency
54
GPT-5.4 ProOpenAI · score 60.89
74 tok/s151.79s latency
55
GPT-5.2OpenAI · score 58.43
73 tok/s130.34s latency
58
Grok 4xAI · score 60.42
54 tok/s15.6s latency
59
GLM-4.5Z.AI · score 57.56
51 tok/s1.45s latency
63
Kimi K2.5Moonshot AI · score 59.66
45 tok/s2.38s latency
66
Kimi K2Moonshot AI · score 27.19
43 tok/s1.51s latency
71
Phi-4Microsoft · score 22.69
35 tok/s2.02s latency
72
GPT-4o miniOpenAI · score 37.87
33 tok/s3.16s latency
73
Gemma 3 27BGoogle · score 41.57
31 tok/s2.04s latency
74
GPT-4 TurboOpenAI · score 27.44
30 tok/s2.84s latency
78
o3-proOpenAI · score 48.33
27 tok/s84.93s latency

Speed data sourced from Artificial Analysis. Metrics reflect median performance across providers. Reasoning models typically show higher first-answer latency due to chain-of-thought processing.

Frequently Asked Questions

What does tokens per second mean for LLMs?

Tokens per second (tok/s) measures how fast an LLM generates output text. Higher is better. A model at 200 tok/s produces roughly 150 words per second — fast enough for real-time streaming. Models below 50 tok/s may feel sluggish in interactive applications.

What does the latency column measure?

Latency here is the time from sending a request to receiving the first token of the answer (Artificial Analysis’s "first answer chunk" metric). Lower is better; under 1 second feels instant in chat. For reasoning models this includes the entire thinking phase, so it can reach 10–150s — it is end-to-end response latency, not raw time-to-first-token of the stream.

Which LLM is the fastest?

Currently, Mercury 2 by Inception is the fastest at 789 tokens/second. The fastest model scoring above 70 overall is GPT-5.4 at 74 tok/s.

Why are reasoning models slower?

Reasoning models (like o3, GPT-5, Gemini Deep Think) use chain-of-thought processing — they generate internal "thinking" tokens before producing the final answer. This adds significant first-answer latency (often 10-150 seconds) but can dramatically improve accuracy on complex tasks. The output speed (tok/s) once generation starts is usually comparable to standard models.

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.