AudioLab Leaderboards

Pairwise preference rankings and independent benchmarks for leading speech models.

How rankings work

Valid votes

67

Models

6

Method

Bradley–Terry MLE

Speech to Text rankings

Models are ordered by rating. Models are marked preliminary until 200 valid battles; disabled models are hidden. Newly added models appear unranked at the bottom until they record their first battles.

Controls only for which response was heard first.

Preliminary models have fewer than the required valid battles or a weakly connected matchup graph. Treat their rank and confidence interval as directional.

Performance by language

Win rate per language. Cells with too few battles are left blank.

ModelEnglish
Mistral Voxtral Small52%
OpenAI GPT-4o Audio52%
Alibaba Qwen3-Omni Flash36%