
ElevenLabs has released AA-WER v2.0, an updated speech-to-text accuracy benchmark that introduces a new proprietary dataset focused specifically on speech directed at voice agents. The update aims to provide a more rigorous and realistic measure of transcription accuracy for conversational AI systems.
The centerpiece of v2.0 is AA-AgentTalk, a held-out dataset comprising 469 samples and roughly 250 minutes of audio captured from voice agent interactions. Unlike public benchmarks that models can train against, AA-AgentTalk remains private to prevent overfitting. The dataset spans voice agent and call center interactions, AI agent conversations, industry jargon, meetings, consumer scenarios, and media content across 17 accent groups, 8 speaking styles, and varied recording devices and environments. AA-AgentTalk now accounts for 50% of the overall benchmark weighting.
Cleaning Up Ground Truth Errors
A significant portion of the v2.0 effort involved correcting errors in existing public datasets. ElevenLabs identified instances where reference transcripts in VoxPopuli and Earnings22 did not accurately reflect what speakers actually said. These discrepancies unfairly penalized models that transcribed the audio correctly.
The company manually reviewed and created cleaned versions of both datasets, now available on Hugging Face as VoxPopuli-Cleaned-AA and Earnings22-Cleaned-AA. Each cleaned dataset accounts for 25% of the new benchmark weighting. Meanwhile, the AMI-SDM dataset was removed entirely due to extensive transcript errors, including heavily overlapping speech that required too many judgment calls to correct reliably.
Better Text Normalization
ElevenLabs also overhauled text normalization to ensure models are evaluated on genuine transcription accuracy rather than formatting quirks. Building on OpenAI's Whisper normalizer package, the custom normalizer addresses digit splitting (preventing mismatches like "1405 553 272" versus "1405553272"), preserves leading zeros, normalizes spoken symbols such as "+" and "_", strips redundant ":00" in times ("7:00pm" versus "7pm"), adds US and UK English spelling equivalences ("totalled" versus "totaled"), and accepts equivalent spellings for ambiguous proper nouns in the dataset ("Mateo" versus "Matteo").
These changes reduce artificially inflated word error rates caused by surface-level formatting differences rather than true transcription mistakes.

Benchmark Results
Under the new AA-WER v2.0 benchmark, ElevenLabs's Scribe v2 leads with a 2.3% error rate, followed by Google DeepMind's Gemini 3 Pro at 2.9%, Mistral AI's Voxtral Small at 3.0%, Google's Gemini 3 Flash at 3.1%, and ElevenLabs Scribe v1 at 3.2%.
Scribe v2 topped two of the three component datasets, AA-AgentTalk and Earnings22-Cleaned-AA, while Gemini 3 Pro led on VoxPopuli-Cleaned-AA. The results underscore the value of domain-specific benchmarks for evaluating models intended for conversational AI and voice agent use cases.
By introducing a proprietary held-out dataset and correcting ground truth errors in public benchmarks, ElevenLabs is pushing the industry toward more reliable measures of speech-to-text performance. The cleaned datasets and improved normalization methods are now available for other researchers and developers to use, potentially raising the bar for the entire field.
Sources
1 checkedHow we cover tool news: Create With's tool desk drafts these reports with AI from the sources listed above and checks them against those sources before publishing.




