
ElevenLabs released Dubbing v2, a new dubbing model that preserves the performance of original audio across languages rather than translating words alone. The model conditions directly on the source audio to maintain tone, emotion, and delivery, with sync-aware translation logic that aligns phrasing to the original timing.
The announcement, posted on LinkedIn on May 28, positions the model as an automated alternative to professional dubbing, which typically costs hundreds of dollars per minute. Dubbing v2 is available in ElevenCreative and via ElevenProductions, with API access coming soon.
Conditioning on Original Audio
Unlike earlier dubbing models that translate text first and synthesize speech separately, Dubbing v2 conditions directly on the original audio track. This approach allows the model to carry over vocal characteristics such as pitch variation, emotional inflection, and pacing alongside the translated content.
The sync-aware translation system adapts phrasing to align with the original speaker's timing, matching starts, stops, and pauses rather than generating a new speech rhythm. ElevenLabs claims this eliminates the flat, disconnected feel common in machine-translated voiceovers.
The model supports 90+ languages, enabling creators to localize video content, marketing campaigns, and broadcast productions without rebuilding audio pipelines or hiring dubbing studios.
Availability and Use Cases
Dubbing v2 is available immediately through ElevenCreative, the platform's content creation suite, and ElevenProductions, which handles larger-scale commercial projects. API access is scheduled for a future release, allowing developers to integrate the model into custom workflows.
The target use cases span YouTube creators localizing content for international audiences, marketing teams producing region-specific ad variations, and studios generating broadcast-quality dubbed versions without manual voice direction. The automation reduces both cost and production time compared to traditional dubbing workflows.
ElevenLabs has been expanding its audio AI capabilities rapidly. The company launched Music v2 with better vocals and a 50% price cut just two days earlier, and recently introduced Speech Engine to convert chat agents into voice agents with a single prompt.
The new dubbing model arrives as demand for automated localization grows. While professional dubbing remains standard for high-budget film and television, automated systems are gaining traction in digital video, e-learning, and marketing where production budgets are tighter and turnaround times are shorter.
ElevenLabs has not released technical details on model architecture, training data, or performance benchmarks against existing dubbing systems. Pricing for Dubbing v2 through ElevenCreative and ElevenProductions was not disclosed in the announcement.
Sources
1 checkedHow we cover tool news: Create With's tool desk drafts these reports with AI from the sources listed above and checks them against those sources before publishing.




