Microsoft AI 发布流式语音模型 MAI-Transcribe-2-Streaming,AA-WER Streaming 排名第一
R to @ArtificialAnlys: On Final Transcript, MAI-Transcribe-2-Streaming achieves a 2.5% WER at 0.13s after end of speech, ranking #1 of 38 models on AA-WER Streaming. It is more accurate than Grok Voice Transcribe 2.0 at 2.7% WER and 0.49s, Muse Voice Transcribe at 3.1% WER and 0.16s, Cartesia Ink Preview with external endpoints at 3.1% WER and 0.11s, and ElevenLabs Scribe v2 Realtime at 3.6% WER and 0.14s, and faster than all of them except Cartesia Ink Preview. It sits at the low-error end of the Pareto frontier for accuracy against time to final transcript.
Microsoft AI 发布流式语音转文字模型 MAI-Transcribe-2-Streaming,在 AA-WER Streaming 的 Final Transcript 与 First Partial Transcript 两项上均排名第一,最终转写 WER 为 2.5%、语音结束后 0.13s 返回。
来源:X:Artificial Analysis (@ArtificialAnlys) · x.lingyaoai.com