ElevenLabs has unveiled its latest speech model, Eleven v4, which significantly improves the expressiveness and consistency of AI-generated voices. According to The Decoder, this new model can accurately interpret cues such as laughter and whispering, maintaining voice consistency across lengthy audio productions like audiobooks.
The Turbo variant of Eleven v4 is designed for real-time applications, capable of starting speech within 150 milliseconds, making it ideal for voice agents. The model's performance places it at the top of Artificial Analysis' Voice Arena leaderboard, surpassing competitors including Cartesia and Google's Gemini.
As Japan continues to integrate AI-driven voice technologies in sectors like customer service and media, advancements like Eleven v4 could accelerate adoption and enhance user experiences in local markets.
