New architecture, sub-second Turbo path and 10-second clones go live in the API, ElevenAgents and ElevenCreative
ElevenLabs on Monday released Eleven v4 and a faster sibling, Eleven v4 Turbo, two text-to-speech models the company says were built to carry tone, not just words.
The post is by co-founder and CEO Mati Staniszewski and Piotr Dabkowski. Both models are live in ElevenAgents, ElevenCreative and the ElevenAPI. A free account is enough to generate with either one.
The problem they named is familiar to anyone who has sat through a support bot. The same sentence does different work in different rooms. “I need you to stay calm” is not the same line in a clinic and in a drop-into-combat game. Older systems could read the words. They often missed the weather around them.
ElevenLabs said v4 is meant to read tone, pacing, emotion, character and context, then keep the speaker’s identity intact. Multi-speaker scenes are supposed to answer what was just said, not splice isolated takes. Ranked No. 1 on Artificial Analysis’s Provider Voice Arena leaderboard for September 2026, the company also said about 75% of listeners preferred v4 in blind head-to-head tests against Cartesia Sonic 3.6, Inworld TTS-2, Google Gemini 3.8 Flash-Lite TTS and Google Gemini 3.8 Flash TTS.
The stack under it is new, not a polish pass. v4 Turbo is the agent cut. ElevenLabs put median inference latency at about 100 milliseconds, which it said is quicker than the average pause between two people talking. Separately, it put median time to first speech at about 150 milliseconds. Those are different clocks. Both matter if a caller is already angry.
Direction is in the prompt. Users can write how a line should land in ordinary language, then drop inline tags such as [laughs], [said angrily in French accent], [light rain] or [phone buzzing]. The company said v4 follows those tags more tightly than earlier models. International Phonetic Alphabet phonemes are improved for custom pronunciations. Full tag syntax sits in the developer docs.
Voice cloning is the other lever. Instant Voice Clones now take about 10 seconds of source audio, according to the post. Professional Voice Clones are supported on v4 for higher-fidelity jobs. Speaker identity is supposed to hold across generations, dialogue, narration and regenerated lines, which is the difference between an audiobook that sounds like one person and one that drifts by chapter. Request stitching — chaining generations for long pieces — is called out as more reliable in Studio and the Reader app.
Language count is more than 90. A voice recorded in one language is supposed to speak the others fluently, take a native accent and keep the original identity. ElevenLabs said that accent no longer slides back toward the source over a long generation. Dubbing and localization are the obvious buyers: one brand voice, many markets.
Turbo is wired to ElevenAgents on purpose. The company argued that shops stitching a voice model from one vendor and an agent stack from another cannot tune the pair as a single system. Its research and engineering teams trained Turbo and ElevenAgents together. Healthcare and games are the examples it used: a calm agent that can say drug names, or a fast one that can talk slang without going flat.
What is not in the post is a price sheet, a rate-limit table or an independent audit of the 75% preference figure. Artificial Analysis is named. The blind test is ElevenLabs’ own, dated September 2026, against four listed rivals. Buyers who need a bake-off still have to run one.
For an executive, the decision is whether voice is a cost line or a brand surface. If it is the latter, v4 is the company’s bid to stop trading warmth for speed. For a developer, the surfaces are the API, the tag syntax, IPA, 10-second instant clones and PVC. Turbo is the path if the first syllable has to leave the stack in a couple of hundred milliseconds.
The models are up. The next test is a real queue, a real chapter, a real localized ad — not a leaderboard row.

