AudioLab Leaderboards

Pairwise preference rankings and independent benchmarks for leading speech models.

How rankings work

Valid votes

190

Models

6

Method

Bradley–Terry MLE

Speech to Speech rankings

Models are ordered by rating. Models are marked preliminary until 200 valid battles; disabled models are hidden. Newly added models appear unranked at the bottom until they record their first battles.

Controls only for which response was heard first.

Preliminary models have fewer than the required valid battles or a weakly connected matchup graph. Treat their rank and confidence interval as directional.

Performance by language

Win rate per language. Cells with too few battles are left blank.

ModelEnglish
Google Gemini 2.5 Flash Native Audio66%
OpenAI GPT Realtime54%
OpenAI GPT-4o Audio50%
OpenAI GPT Realtime 1.557%
xAI Grok Voice38%
Alibaba Qwen3-Omni Flash31%