Alibaba’s Qwen Audio 3.0 TTS Plus has secured the top position on Artificial Analysis’ Speech Arena leaderboard for text-to-speech (TTS) models, according to The Decoder. The model supports 16 languages and allows users to customize speaking styles using natural language or tags such as [angry], enhancing its versatility in voice synthesis.
Despite its advanced features and broad language coverage, Qwen Audio 3.0 TTS Plus processes speech at 16 characters per second, which is slower compared to rival models Sonic 3.5 and Simba 3.2, as reported by The Decoder. This trade-off highlights Alibaba’s focus on quality and expressiveness over raw speed.
For Japanese markets, where nuanced voice technology plays a growing role in customer service and AI applications, Alibaba’s TTS advancements underscore increasing competition in multilingual speech synthesis solutions across Asia.
