跳到正文
原文
rohanpaul_ai· @rohanpaul_ai · X·· 16 小时前AI 评分44

弱语言ASR:更多数据胜过大模型

AI 导读

对于弱语言的 ASR(自动语音识别),更多训练数据往往胜过更大的模型。 @voicearena_ai 的孟加拉语结果就证明了这一点,Monsoon 是其 ASS 语料库。 在 Monsoon 上微调的 Whisper Medium,在孟加拉语 FLEURS 上的 LLM 词错误率从 85.27% 降至 7.65%。 Medium 是 769M 参数的 Whisper。足够小,可以低成本部署,在许多场景下也足够小,可以靠近用户运行。

正文 · 原文

For ASR (automatic speech recognition) on weak languages, more training data tends to beat a bigger model.

@voicearena_ai’s Bengali result for Monsoon, its ASR corpus, is shows it.

Whisper Medium fine-tuned on Monsoon went from 85.27% to 7.65% LLM word error rate on Bengali FLEURS.

Medium is the 769M-parameter Whisper. Small enough to serve cheaply, and in many settings small enough to run close to the user.

来源:rohanpaul_ai · x.com