跳到正文
原文
rohanpaul_ai· @rohanpaul_ai · X·· 13 小时前AI 评分41

smallest AI 在 VoiceArena 语音基准登顶

AI 导读

smallest AI 在 VoiceArena 的说话人分离(diarization)+ ASR 基准上夺得第一。该基准让模型听远场房间音频,但用每位说话人身上单独佩戴的麦克风所建标签来评分,实现"真实输入 + 精确 ground truth"的组合。作者指出许多 diarization 基准要么用干净音频、要么标签粗糙,VoiceArena 两者都避免,因此这个第一含金量高。

正文 · 原文

Congrats @smallest_AI on taking #1 for diarization + ASR on @voicearena_ai.

Its a very effective real-world kind of benchmark:

> models hear far-field room audio (i.e. microphone is some distance away from the people speaking), but they’re scored against labels built from a mic on every speaker.

> Realistic input, accurate ground truth. That combination is rare.

Many diarization benchmarks either use clean audio or have messy labels.

@voicearena_ai avoids both: models get far-field recordings of real in-person conversations, and the ground truth comes from individual mics on each speaker. #1 really means something.

Congrats @smallest_AI.

来源:rohanpaul_ai · x.com