Alibaba's Qwen Audio 3.0 TTS Plus tops the competition in the text-to-speech rankings
By Matthias Bastian

AI 摘要
Alibaba's Qwen Audio 3.0 TTS Plus tops Artificial Analysis' Speech Arena leaderboard. The model supports 16 languages and lets users control speaking style with natural language or tags like [angry]. At 16 characters per second, though, it's far slower than rivals Sonic 3.5 and Simba 3.2. The arti
原文正文
Alibaba's Qwen Audio 3.0 TTS Plus tops the competition in the text-to-speech rankings
Alibaba's new text-to-speech model, Qwen-Audio-3.0-TTS-Plus, leads Artificial Analysis' Speech Arena leaderboard for provider voices. With an Elo score of 1,236, it sits just ahead of Simba 3.2 (1,234). Gemini 3.1 Flash TTS (1,214) and Sonic 3.5 (1,207) follow behind.
The model comes in two versions. Flash is built for real-time interaction with about 300 milliseconds of latency, while Plus targets high-quality speech output. It supports 16 languages, including less commonly covered ones like Tagalog, Malay, Thai, and Vietnamese, along with several Chinese dialects. Users can steer the speaking style with natural language or add nonverbal cues using tags like "[angry]" or "[giggles]." Alibaba also says the model handles noisy or echo-heavy reference recordings better than previous versions when cloning voices.
Speed is a weak spot: At 16 characters per second, it trails Sonic 3.5 (120) and Simba 3.2 (30.2) by a wide margin. Pricing lands at $27.60 per million characters through Alibaba Cloud Model Studio. A collection of audio samples is available here.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Subscribe now