Alibaba releases Qwen Audio 3.1 models
Alibaba's AI team Qwen released Qwen-Audio-3.1, a lineup of five models for speech recognition, text-to-speech, and real-time interaction. The automatic speech recognition model improves multilingual and dialect recognition while cleaning up filler words and repetitions. ASR-Next adds multi-speaker identification with timestamps and detects emotions, ambient sounds, and machine noise. Text-to-speech handles multilingual synthesis, allowing users to control emotion, speed, and style through simple text prompts.
TTS-Next pairs a language model with a diffusion approach to generate voice, sound effects, and background audio in a single pass. The real-time model supports simultaneous speaking and listening with instant interruption and adapts its response when it detects a low mood. Alibaba is also slashing prices across the lineup. Text-to-speech drops about 70 percent, the real-time model drops roughly 85 percent, and speech recognition drops up to 95 percent.