Alibaba Releases Qwen-Audio-3.1: Adds ASR-Next, TTS-Next, and Slashes Prices Up to 95%

Alibaba has unveiled the Qwen-Audio-3.1 stack, introducing ASR-Next and TTS-Next alongside major API price cuts: up to 95% off ASR, 85% off Realtime, and 70% of

tau · September 23, 2026

#Qwen #Alibaba #QwenAudio #VoiceAI #TTS #ASR #Realtime

Alibaba Releases Qwen-Audio-3.1: Adds ASR-Next, TTS-Next, and Slashes Prices Up to 95%

On September 23, 2026, Alibaba's Qwen team officially unveiled 'Qwen-Audio-3.1', a comprehensive audio model suite spanning speech understanding, generation, real-time interaction, and multimodal audio creation. The release overhauls the existing ASR (speech recognition), TTS (text-to-speech), and Realtime models while introducing two new architectures—TTS-Next for unified audio generation and ASR-Next for advanced audio understanding—establishing a complete five-model audio stack.

Official announcement graphic for Alibaba Qwen-Audio-3.1 audio model suite

Image source: @Alibaba_Qwen

Alongside the expanded model line, Alibaba introduced aggressive price cuts across its audio portfolio on the Qwen cloud platform. The company reduced API rates by up to 95% for ASR, approximately 85% for Realtime duplex calls, and around 70% for TTS speech synthesis, lowering infrastructure cost barriers for voice AI deployments.

Five-Model Audio Architecture: Upgraded Foundations with ASR-Next and TTS-Next

The Qwen-Audio-3.1 release structures its capabilities into four core pillars—understanding, generation, interaction, and creation—distributed across five specialized models.

At the base recognition layer, the standard ASR model received significant multilingual and regional dialect improvements. It now incorporates native polishing that automatically detects and strips conversational filler words and redundant phrasing, producing cleaner, more coherent written transcripts.

The newly introduced 'ASR-Next' moves beyond standard speech-to-text with a next-generation audio understanding architecture. It delivers multi-speaker recognition complete with speaker diarization labels, precise timestamps, and aligned transcripts. Furthermore, ASR-Next extends acoustic perception to identify ambient environmental sounds, mechanical noises, and emotional states, powering sound captioning, acoustic event localization, and rich audio question-answering and reasoning.

On the generation side, responsibilities are split between the standard TTS engine and the new TTS-Next. The upgraded TTS model enhances cross-lingual voice cloning and multilingual dialect synthesis, allowing developers to steer emotional tone, cadence, and speaking style via natural language instructions. The companion 'TTS-Next' model introduces a unified framework combining language modeling (LM) with diffusion architectures. Designed for full audio production, TTS-Next synthesizes human speech, background Foley sound effects, and ambient acoustic environments in a single generation pass, outputting 48kHz high-fidelity multi-character audio suited for audiobooks, podcasts, game assets, and commercial advertising.

Full-Duplex Real-Time Dialogue: Multi-Teacher Distillation and Emotion-Adaptive Interaction

For conversational applications, 'Qwen-Audio-3.1-Realtime' enables full-duplex voice interaction, allowing users and the assistant to speak and listen simultaneously just like in a natural telephone conversation. Moving beyond rigid turn-taking, the system supports seamless anytime interruption, letting users cut in or redirect conversational flow without friction.

Architecturally, Qwen-Audio-3.1-Realtime is built upon a multi-teacher distillation framework designed to maintain low latency while supporting sophisticated interactive capabilities:

  • Emotion Understanding and Adaptive Cadence: The model senses vocal tone and subdued mood; when it detects emotional changes in the user, it automatically moderates its speaking rate and responds with an empathetic delivery.
  • Real-Time Language Switching: It dynamically tracks and shifts languages mid-conversation without dropping interactive context.
  • In-Dialogue Tool Calling: The assistant can invoke tools and external functions during active voice sessions, enabling backend workflows while sustaining continuous audio interaction.

Up to 95% API Price Reductions and Cloud Availability

Beyond architectural enhancements, the immediate practical impact of the Qwen-Audio-3.1 release is a drastic price reduction across Alibaba's audio API portfolio:

  • ASR (Speech Recognition): Discounted by up to 95%
  • Realtime (Duplex Interaction): Discounted by approximately 85%
  • TTS (Speech Synthesis): Discounted by approximately 70%

In terms of deployment status, the base Qwen-Audio-3.1-ASR and Qwen-Audio-3.1-Realtime APIs are immediately accessible on Alibaba's Qwen cloud platform (@qwen_cloud).

However, availability remains phased across the broader lineup. The comprehensive ASR-Next API is designated as 'coming soon', with TTS-Next cloud endpoints expected in a subsequent release. In response to developer inquiries regarding downloadable model weights, Alibaba has not confirmed an open-source timeline at launch, focusing initially on hosted cloud API delivery.

Sources