NNewsGPT ← Home
CN

Alibaba Launches Qwen-Audio-3.0-TTS, a New Large-Scale Speech Synthesis Model

CN16 hr ago

On July 20th, Alibaba announced the release of its new large-scale speech synthesis model, Qwen-Audio-3.0-TTS. This advanced model features systematic improvements in areas such as fine-grained label control, freestyle instruction following, multilingual and dialect coverage, and robustness to complex acoustic conditions. The model is available in two versions: a Flash version optimized for real-time interaction with a first-packet latency of around 300 milliseconds, and a Plus version designed for high-quality generation. Alibaba aims for its synthesized speech to move beyond simply speaking to truly expressing meaning. Demonstrating its capabilities, the Plus version of Qwen-Audio-3.0-TTS has achieved the top rank on the global third-party Artificial Analysis leaderboard. Furthermore, its preview version, Fun-Realtime-TTS, had previously topped the same leaderboard one month prior to this announcement.

AI Analysis

Alibaba's introduction of Qwen-Audio-3.0-TTS signifies a competitive advancement in the rapidly evolving field of generative AI for speech. The focus on granular control, instruction adherence, and broad linguistic coverage addresses key limitations in current text-to-speech technologies, moving towards more nuanced and expressive vocal output. By offering both real-time and high-fidelity versions, Alibaba targets diverse market needs, from interactive applications to professional audio production. The model's performance on third-party benchmarks suggests a strong technical foundation, potentially influencing future industry standards and user expectations for synthetic voice quality and versatility. This development underscores the ongoing race among major tech players to develop and deploy sophisticated AI models across various modalities.

AI-generated to prompt reflection — not editorial opinion, not advice, not a statement of fact. How this works.

Compiled by NewsGPT from 36Kr (CN). Read the original for full details.