ElevenLabs launches v4 voice generation models

Written by

Picture of فريقنا

فريقنا

Communications Consultant

ElevenLabs has launched its fourth-generation and Turbo text-to-speech models featuring a new software architecture that supports over 90 languages. The new models enable voice cloning in just 10 seconds with emotional performance direction.

ElevenLabs has launched its fourth-generation and Turbo text-to-speech models featuring a new software architecture that supports over 90 languages. The new models enable voice cloning in just 10 seconds with emotional performance direction.

New software architecture and expanded language support

Leading voice AI company ElevenLabs announced on Monday the launch of its new voice generation models, Eleven v4 and the ultra-fast version Eleven v4 Turbo. Built on an entirely new neural architecture, the company states they deliver the highest levels of vocal expression and control over tone and performance to date. The new models are immediately available to developers and the public via the company’s API and various consumer products, including the free usage tier.

The new models expand language coverage to support more than 90 global languages, compared to around 70 supported by the previous version, alongside fundamental improvements in pronunciation accuracy and local accent simulation. The company noted that the largest improvements in vocal quality were clearly visible in Japanese, Brazilian Portuguese, Mandarin Chinese, and Cantonese, giving creators unprecedented natural pronunciation tools that eliminate robotic monotony.

Lightning-fast voice cloning and embedded performance direction

The fourth generation introduces a revolutionary feature: the ability to clone any human voice with high precision using a reference audio clip lasting just 10 seconds, representing a massive reduction in the audio data requirements needed for training. The company also added support for embedded text-based voice direction tags, allowing users to include acting instructions and natural language directly within the written text—such as inserting expressions for laughter, whispering, and slamming doors—for the model to translate vocally with striking spontaneity.

This release marks the return of the “professional voice cloning” feature, which was absent from the v3 model released last year. The models support generating long texts of up to 10,000 tokens per batch, equipped with an intelligent “context stitching” feature designed to maintain consistent delivery tone, voice rhythm, and continuity across lengthy audio projects, such as audiobooks and long novels, without any break in style.

Ultra-fast Turbo model for phone conversational agents

The Turbo version tailored for the fourth generation has been optimized specifically for real-time conversational AI applications, supporting bidirectional audio streaming that begins returning and voicing output to the listener even before the text sentence generation is complete within the system. ElevenLabs recorded a striking average time-to-first-audio of approximately 150 milliseconds in the company’s internal benchmark tests, compared to around 262 milliseconds for the Cartesia Sonic 3.6 model and 814 milliseconds for the OpenAI GPT-4o mini audio model.

According to a report published by TechCrunch, the model features advanced behavioral capabilities to manage difficult conversational scenarios during automated calls; it can handle verbal interruptions flexibly, deal with complaint escalations, and handle call parking and customer hold states. These advanced features are specifically designed to meet the needs of call centers and customer service departments in major enterprises seeking reliable automation that mimics human behavior.

Financial revenue surge and public offering aspirations

ElevenLabs has achieved rapid and remarkable growth in its infrastructure and business since raising $500 million in funding from Sequoia Capital earlier this year at an $11 billion valuation. The company’s annualized run-rate revenue jumped from approximately $330 million at the beginning of the year to surpass the $600 million mark, with over 55% of this revenue flowing directly from enterprise and large corporate customers, while its employee headcount swelled past 800 across India, Europe, and Brazil.

Strong rumors are circulating in financial markets regarding an upcoming funding round that could push the company’s valuation to $22 billion. Mati Staniszewski, co-founder and CEO, told TechCrunch that the company aims to execute an initial public offering (IPO) of its shares “over the coming years” without committing to a strict timeline. These steps come amid fierce competition in the generative voice sector with startups like Cartesia and Deepgram, alongside competitive projects from Google and OpenAI.

FAQs

Question: How many languages does the fourth generation of ElevenLabs models support?
Answer: The model supports over 90 global languages with notable improvements in East Asian and European languages.

Question: What is the minimum duration of the audio recording required for voice cloning?
Answer: The new technology requires a reference audio recording of just 10 seconds to clone voice identity with high precision.

Question: What speed does the Turbo model achieve in starting interactive speech?
Answer: The model records an average time-to-first-speech of 150 milliseconds, making it ideal for conversational agents in call centers.

شارك هذا الموضوع:

شارك هذا الموضوع:

اترك رد

الفئات

المنشورات الأخيرة

Discover more from Buzzinga

Subscribe now to keep reading and get access to the full archive.

Continue reading