ElevenLabs releases Eleven v4 speech model - more than 90 languages and voice cloning from 10 seconds of audio
ElevenLabs released Eleven v4, a new text-to-speech model, along with a faster v4 Turbo variant for voice agents, whose time to first audio is about 150 milliseconds (median), according to the company. Both models support more than 90 languages, up from 70 in the previous version, and 10 seconds of audio is enough to clone a voice. The models are already available in ElevenLabs services and via the API.