Eleven v4 takes the English TTS lead at a premium
At publication, Artificial Analysis’ Provider Voice Arena lists Eleven v4 first among English text-to-speech models. It ranks ahead of Cartesia’s Sonic 3.6 and Google’s Gemini 3.8 Flash TTS while setting the benchmark’s highest composite score for pronunciation robustness.
ElevenLabs launched two variants for different workloads. The flagship v4 targets narration and character performances, while Turbo prioritizes low-latency voice agents. TechCrunch reports that both add expression controls and support more than 90 languages, up from roughly 70 in v3.
Native voices put v4 ahead
The Provider Voice Arena evaluates each model with its own voices and default configuration. Its Elo score summarizes head-to-head listener preferences, with a higher score indicating that evaluators preferred the model more often.
Eleven v4’s score draws on 1,674 evaluation samples. The model also ranks first in each of the arena’s four content categories: customer service, assistants, knowledge sharing, and entertainment.
The Controlled Voice board uses the same cloned voice across systems, reducing the influence of each provider’s voice catalog. Eleven v4 ranks second there with 1,157 Elo, behind Alibaba’s Qwen-Audio-3.1-TTS-Plus at 1,178. The result suggests that ElevenLabs’ native voices and default configuration contribute to its lead on the main board.
Elo scores are relative rankings rather than quality percentages, and they can move as evaluators add votes. A 21-point gap on the Controlled Voice board should therefore inform direct testing, not replace it.
Pronunciation reaches 91.7%
Speech models frequently misread abbreviations, technical terms, context-dependent words, and character sequences such as order numbers. Artificial Analysis tests those cases by asking human reviewers to compare generated clips with agreed pronunciations.
The 91.7% composite is the highest Artificial Analysis has measured for a TTS model. Eleven v4 is also the only model above 93% in three of the four categories.
Exact character sequences remain its weakest category at 78.8%. Applications that read confirmation numbers, medical terms, account identifiers, or product SKUs still need production testing and a fallback for high-risk strings.
Throughput rises alongside price
Artificial Analysis measured Eleven v4 at 73.4 generated characters per second, up from 42.5 for v3. That figure measures synthesis throughput; it does not include network delay, language-model response time, or client-side audio playback.
Eleven v4 costs about 1.6 times as much as Sonic 3.6 and 4.9 times as much as Gemini 3.8 Flash TTS at listed standard rates. Pricing may vary by plan, volume agreement, and included credits.
A two-week launch promotion reduces v4 to $22 per million characters and Turbo to $11 per million. Eligible Creator+ subscriptions can also use v4 within their monthly credit limits without an additional model surcharge. The standard $80 rate becomes the relevant long-term figure after the promotion expires.
Cloning changes require migration
ElevenLabs rebuilt its voice-cloning stack for v4. Instant cloning can create a voice from a 10-second recording, while professional cloning returns after being unavailable in v3. The company says v4 instant clones outperform professional clones produced with Multilingual v2, although the public arena results do not independently test that specific claim.
Professional and instant clones created before v4 require retraining for effective use with the new model. Applications already calling ElevenLabs’ conversion API can select v4 by changing the model identifier:
const audio = await elevenlabs.textToSpeech.convert(
's3TPKV1kjDlVtZbl4Ksh',
{
text: 'The first move is what sets everything in motion.',
modelId: 'eleven_v4'
}
);Clone retraining remains a separate migration step from the SDK change. Teams should compare old and retrained voices for identity, pacing, pronunciation, and consistency before moving production traffic.
Turbo reports a median time to first speech of 150 milliseconds and supports bidirectional streaming. Time to first speech measures the interval between a synthesis request and the first returned audio chunk. End-to-end agent latency also depends on network conditions, language-model generation, buffering, and playback.
The English lead has limits
ElevenLabs’ benchmark lead applies specifically to the English Provider Voice Arena. Cartesia’s Sonic 3.6 still leads eight of the nine non-English language boards tracked by Artificial Analysis, despite v4’s support for more than 90 languages.
Competition also includes Deepgram, Fish Audio, Boson, WellSaid Labs, Google, OpenAI, and self-hosted models. On the Controlled Voice board, Qwen-Audio-3.1-TTS-Plus retains the lead, giving teams with a fixed cloned voice another strong candidate to evaluate.
ElevenLabs reports that enterprise customers generate more than 55% of its revenue. Its annualized recurring revenue has exceeded $600 million, up from roughly $330 million at the start of 2026. The company also raised $500 million in a Sequoia-led round at an $11 billion valuation.
Match the model to the workload
- English narration and character work: Eleven v4 leads the native-voice preference board and offers expanded expression controls, with a standard price of $80 per million characters.
- Real-time voice agents: Turbo targets faster conversational loops through streaming and a reported 150-millisecond median time to first speech.
- Fixed cloned voices: Qwen-Audio-3.1-TTS-Plus leads the Controlled Voice board, making same-voice comparisons especially relevant.
- High-volume, cost-sensitive synthesis: Gemini 3.8 Flash TTS carries a substantially lower listed character price.
- Non-English speech: Sonic 3.6 leads most of the language-specific boards tracked by Artificial Analysis.
- Self-hosted deployment: Open-weight systems such as Breeze TTS 2 exchange per-character API fees for infrastructure and operational work.
Production evaluations should use representative scripts, voices, languages, and identifiers from the intended application. Eleven v4’s leaderboard position makes it a leading candidate for English workloads, while its price, clone-migration requirement, and weaker exact-sequence score remain material deployment considerations.