评测五种主流说话人分离模型,揭示错误根源与性能差距。
Benchmarking Diarization Models
- 在多语言音频上对比五款先进模型的分离效果。
- 最佳模型PyannoteAI达11.2%错误率,开源模型DiariZen为13.3%。
- 高人数对话中漏语音是主因,易引发说话人混淆。
说话人分离旨在根据说话人身份对音频进行分段,回答多说话人对话中“谁在何时说话”的问题。尽管该任务对下游应用至关重要,仍属未解难题,其错误会传递至下游系统并引发广泛故障。为此,我们在涵盖多种语言和声学条件的四个数据集上,评估了五种最先进的分离模型。数据集共包含196.6小时多语言音频,涵盖英语、中文、德语、日语和西班牙语。结果表明,PyannoteAI表现最优,错误率为11.2%;而DiariZen提供具有竞争力的开源替代方案,错误率为13.3%。分析失败案例发现,主要错误源于遗漏语音段,随后引发说话人混淆,尤其在高说话人数量场景下更为显著。
原文摘要 · Abstract (English)
Speaker diarization is the task of partitioning audio into segments according to speaker identity, answering the question of "who spoke when" in multi-speaker conversation recordings. While diarization is an essential task for many downstream applications, it remains an unsolved problem. Errors in diarization propagate to downstream systems and cause wide-ranging failures. To this end, we examine exact failure modes by evaluating five state-of-the-art diarization models, across four diarization datasets spanning multiple languages and acoustic conditions. The evaluation datasets consist of 196.6 hours of multilingual audio, including English, Mandarin, German, Japanese, and Spanish. Overall, we find that PyannoteAI achieves the best performance at 11.2% DER, while DiariZen provides a competitive open-source alternative at 13.3% DER. When analyzing failure cases, we find that the primary cause of diarization errors stem from missed speech segments followed by speaker confusion, especially in high-speaker count settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。