Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief
Live source snapshot

Audio Realism Benchmark

Blind listening comparisons measure how human text-to-speech output sounds across phone-agent, conversational, and explainer speech.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
500 held-out American-English prompts; female and male voices; human recordings included in the comparison pool
Primary metric
Bradley–Terry Elo-style rating, win rate, and uncertainty
Owner
Design Arena
Available evidence
Machine-readable public leaderboard

What it measures

Voice experience
Is the exchange natural, robust, and responsive?

How the protocol works

  1. 01

    Vetted native speakers hear two clips generated from the same prompt and select the one that sounds more human.

  2. 02

    The benchmark uses 500 private, held-out prompts split across phone-agent, conversational, and explainer speech.

  3. 03

    One female and one male American-English voice are selected for each model, and comparisons use the same gender.

  4. 04

    Scores are fitted with a Bradley–Terry model and presented as an Elo-style rating; real human recordings remain in the comparison pool.

Available results

Protocol details
RankModelElo-style ratingBT uncertaintyWin rateBattlesAvg. generation
1Bland Speech v3
Bland AI
1356±22.682.1%4031.62s
2MAI-Voice-2
Microsoft
1223±18.969.0%4091.49s
3Grok TTS
xAI
1168±18.362.3%4032.26s
4Gemini 2.5 Pro TTS Preview
Google
1111±17.555.4%4196.91s
5MiniMax Speech-02 HD
MiniMax
1098±17.852.8%4113.27s
6Cartesia Sonic 3.5
Cartesia
1089±17.550.9%4241.52s
7Gemini 3.1 Flash TTS Preview
Google
1040±17.741.7%4174.91s
8GPT-4o mini TTS
OpenAI
1014±18.142.4%3961.99s
9Gemini 2.5 Flash TTS Preview
Google
992±18.138.3%4154.98s
10Lightning v3.1 Pro
Smallest AI
955±18.733.9%4071.90s
11ElevenLabs Eleven v3
ElevenLabs
949±18.830.8%4093.25s
12Inworld TTS-1.5 Max
Inworld
771±23.314.8%4134.94s

Interpretation limit

The leaderboard measures perceived human-likeness for the selected American-English voices and prompt mix. It does not measure factual accuracy, task completion, full-duplex interaction, or multilingual quality.