LLM Speed & Latency Comparison
Compare inference speed across every major AI model. Tokens/sec measures output generation speed. Latency measures time to the first answer token — for reasoning models this includes thinking time, so it reflects end-to-end response latency rather than raw TTFT.
Speed data from Artificial Analysis. Last updated: 2026-07-21. Median tokens/s, Latency first answer chunk (s).
Mercury 2
789 tok/s · Inception
Command A+
0.25s to first answer · Cohere
GPT-5.4
74 tok/s · Score: 74.24
Top 15 — Output Speed (tok/s)
- Mercury 2789 tok/s, Inception, 3.88 seconds first-answer latency, score 51.28
- Nemotron 3 Super 100B367 tok/s, NVIDIA, 0.71 seconds first-answer latency, score 50.08
- GPT-OSS 20B313 tok/s, OpenAI, 0.65 seconds first-answer latency, score 42.74
- Gemini 3.5 Flash284.2 tok/s, Google, 18.55 seconds first-answer latency, score 64.75
- Ministral 3 3B274 tok/s, Mistral, 0.42 seconds first-answer latency, score 18.27
- Command A+272 tok/s, Cohere, 0.25 seconds first-answer latency, score 47.51
- GPT-OSS 120B262 tok/s, OpenAI, 0.79 seconds first-answer latency, score 50.08
- Grok 4.20233 tok/s, xAI, 10.33 seconds first-answer latency, score 54.68
- Gemini 2.5 Flash221 tok/s, Google, 0.5 seconds first-answer latency, score 48.09
- Ling 2.6 Flash209.5 tok/s, InclusionAI, 1.07 seconds first-answer latency, score 43.87
- Grok 4.3209 tok/s, xAI, 12.36 seconds first-answer latency, score 65.1
- Gemini 3.1 Flash-Lite205 tok/s, Google, 7.5 seconds first-answer latency, score 50.83
- GPT-5.4 mini201 tok/s, OpenAI, 3.85 seconds first-answer latency, score 56.77
- GPT-5.4 nano191 tok/s, OpenAI, 3.64 seconds first-answer latency, score 66.79
- Grok 3 Mini190 tok/s, xAI, 0.54 seconds first-answer latency
Bar length is relative to the fastest model in the current filtered view. Exact median speed is shown at right; open a model for its complete evidence profile.
Average Speed by Provider
NVIDIA
260 tok/s avg · 2 models
1.3s avg latency
xAI
166 tok/s avg · 6 models
7s avg latency
154 tok/s avg · 8 models
14.2s avg latency
Mistral
126 tok/s avg · 7 models
0.8s avg latency
OpenAI
121 tok/s avg · 25 models
42.3s avg latency
Meta
93 tok/s avg · 3 models
1.3s avg latency
Z.AI
82 tok/s avg · 5 models
1.3s avg latency
Anthropic
52 tok/s avg · 7 models
3.3s avg latency
DeepSeek
48 tok/s avg · 2 models
2.3s avg latency
MiniMax
46 tok/s avg · 2 models
2.3s avg latency
Moonshot AI
44 tok/s avg · 2 models
1.9s avg latency
78 models match these filters.
Speed data sourced from Artificial Analysis. Metrics reflect median performance across providers. Reasoning models typically show higher first-answer latency due to chain-of-thought processing.
Frequently Asked Questions
What does tokens per second mean for LLMs?
Tokens per second (tok/s) measures how fast an LLM generates output text. Higher is better. A model at 200 tok/s produces roughly 150 words per second — fast enough for real-time streaming. Models below 50 tok/s may feel sluggish in interactive applications.
What does the latency column measure?
Latency here is the time from sending a request to receiving the first token of the answer (Artificial Analysis’s "first answer chunk" metric). Lower is better; under 1 second feels instant in chat. For reasoning models this includes the entire thinking phase, so it can reach 10–150s — it is end-to-end response latency, not raw time-to-first-token of the stream.
Which LLM is the fastest?
Currently, Mercury 2 by Inception is the fastest at 789 tokens/second. The fastest model scoring above 70 overall is GPT-5.4 at 74 tok/s.
Why are reasoning models slower?
Reasoning models (like o3, GPT-5, Gemini Deep Think) use chain-of-thought processing — they generate internal "thinking" tokens before producing the final answer. This adds significant first-answer latency (often 10-150 seconds) but can dramatically improve accuracy on complex tasks. The output speed (tok/s) once generation starts is usually comparable to standard models.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.