Skip to content
All tool news

Tool desk · Updated daily

Tool news·ElevenLabs·

ElevenLabs Unveils Tougher Speech-to-Text Benchmark With Private Agent Dataset.

ElevenLabs releases AA-WER v2.0 benchmark with proprietary AA-AgentTalk dataset focused on voice agent speech. Scribe v2 leads at 2.3% error rate.

CW

Create With tool desk

1 source checked · 2 min read · 3 sections

ShareLinkedIn
ElevenLabs Unveils Tougher Speech-to-Text Benchmark With Private Agent Dataset

ElevenLabs has released AA-WER v2.0, an updated speech-to-text accuracy benchmark that introduces a new proprietary dataset focused specifically on speech directed at voice agents. The update aims to provide a more rigorous and realistic measure of transcription accuracy for conversational AI systems.

The centerpiece of v2.0 is AA-AgentTalk, a held-out dataset comprising 469 samples and roughly 250 minutes of audio captured from voice agent interactions. Unlike public benchmarks that models can train against, AA-AgentTalk remains private to prevent overfitting. The dataset spans voice agent and call center interactions, AI agent conversations, industry jargon, meetings, consumer scenarios, and media content across 17 accent groups, 8 speaking styles, and varied recording devices and environments. AA-AgentTalk now accounts for 50% of the overall benchmark weighting.

Cleaning Up Ground Truth Errors

A significant portion of the v2.0 effort involved correcting errors in existing public datasets. ElevenLabs identified instances where reference transcripts in VoxPopuli and Earnings22 did not accurately reflect what speakers actually said. These discrepancies unfairly penalized models that transcribed the audio correctly.

The company manually reviewed and created cleaned versions of both datasets, now available on Hugging Face as VoxPopuli-Cleaned-AA and Earnings22-Cleaned-AA. Each cleaned dataset accounts for 25% of the new benchmark weighting. Meanwhile, the AMI-SDM dataset was removed entirely due to extensive transcript errors, including heavily overlapping speech that required too many judgment calls to correct reliably.

Better Text Normalization

ElevenLabs also overhauled text normalization to ensure models are evaluated on genuine transcription accuracy rather than formatting quirks. Building on OpenAI's Whisper normalizer package, the custom normalizer addresses digit splitting (preventing mismatches like "1405 553 272" versus "1405553272"), preserves leading zeros, normalizes spoken symbols such as "+" and "_", strips redundant ":00" in times ("7:00pm" versus "7pm"), adds US and UK English spelling equivalences ("totalled" versus "totaled"), and accepts equivalent spellings for ambiguous proper nouns in the dataset ("Mateo" versus "Matteo").

These changes reduce artificially inflated word error rates caused by surface-level formatting differences rather than true transcription mistakes.

ElevenLabs enterprise voice AI deployment options
ElevenLabs enterprise voice AI deployment options

Benchmark Results

Under the new AA-WER v2.0 benchmark, ElevenLabs's Scribe v2 leads with a 2.3% error rate, followed by Google DeepMind's Gemini 3 Pro at 2.9%, Mistral AI's Voxtral Small at 3.0%, Google's Gemini 3 Flash at 3.1%, and ElevenLabs Scribe v1 at 3.2%.

Scribe v2 topped two of the three component datasets, AA-AgentTalk and Earnings22-Cleaned-AA, while Gemini 3 Pro led on VoxPopuli-Cleaned-AA. The results underscore the value of domain-specific benchmarks for evaluating models intended for conversational AI and voice agent use cases.

By introducing a proprietary held-out dataset and correcting ground truth errors in public benchmarks, ElevenLabs is pushing the industry toward more reliable measures of speech-to-text performance. The cleaned datasets and improved normalization methods are now available for other researchers and developers to use, potentially raising the bar for the entire field.

Sources

1 checked

How we cover tool news: Create With's tool desk drafts these reports with AI from the sources listed above and checks them against those sources before publishing.

Worth passing on?

ShareLinkedIn

Go deeper on ElevenLabs

Related reading, watching and going.

Everything on ElevenLabs →

Latest tool news

What else changed this week.

All tool news
MakeDigest

What Make Shipped in Its Latest Update

Make just made scenario design less punishing. An Undo‑Redo feature landed in the scenario editor, giving builders a safety net when they move, link or delete modules.

Zapier

Zapier Moves Agents Into AI by Zapier

AI by Zapier now contains the tool calling, reasoning and autonomous action previously offered through Zapier Agents. Builders can add those capabilities as a single AI step…

The Create With Briefing

Don't watch forty changelogs. Read one email.

Every Tuesday: the tool changes worth knowing, real business use cases, and what's on near you. Free, unsubscribe any time.