阿里发布 Qwen-Audio-3.1 音频模型系列,含 TTS-Next 与 ASR-Next 并大幅降价
⚡ Meet Qwen-Audio-3.1! ASR, TTS & Realtime are fully upgraded, joined by two new models: TTS-Next for audio creation and ASR-Next for audio understanding. Five models, one complete audio stack: understanding, generation, interaction & creation. Plus big price cuts across the lineup: TTS ~70% off, Realtime ~85% off, and ASR up to 95% off. Highlights: 🥳 - ASR: stronger multilingual & dialect recognition, plus native polishing that auto-removes fillers & repetitions for cleaner, more logical transcripts. - ASR-Next: supports multi-speaker ASR with speaker labels, timestamps & aligned transcripts, and understands emotions, ambient & machine sounds for sound captioning, event localization, audio QA & reasoning. - TTS: multilingual & dialect synthesis with natural cross-lingual voice transfer; control emotion, speed & style via simple instructions. - TTS-Next: unified LM + diffusion framework generating voice, sound effects & background audio in one pass for audiobooks, podcasts, games & ad
阿里发布 Qwen-Audio-3.1,ASR、TTS 与 Realtime 全面升级,并新增 TTS-Next 和 ASR-Next 两款模型,共五款模型覆盖理解、生成、交互与创作。
官方一次性给出五款音频模型的能力分工与降价幅度,可据此判断语音链路的选型与成本变化。
来源:X:通义千问 / Qwen (@Alibaba_Qwen) · x.lingyaoai.com