arXiv:2409.15356eess.AScs.LG2024-09

团队在2024年DISPLACE挑战中构建了多语言多人语音与语言分段系统。

TCG CREST System Description for the Second DISPLACE Challenge

  • 采用语音增强、声活动检测与神经嵌入融合策略,基于SpeechBrain实现
  • 语音分段任务较基准提升约7%相对性能,语言分段未超越基准
  • 适合关注多语种多说话人场景下语音分析的开发者参考

本文描述了本团队为2024年第二届DISPLACE挑战赛开发的说话人分段(SD)与语言分段(LD)系统。研究聚焦于多语言、多说话人场景下的第1赛道(SD)与第2赛道(LD)。我们探索了多种语音增强技术、声活动检测(VAD)方法、无监督领域分类以及神经嵌入提取架构,并尝试融合不同嵌入模型。系统基于开源SpeechBrain工具包实现,最终提交方案均采用谱聚类进行分段。在第1赛道中,系统相较挑战基线实现约7%的相对性能提升;第2赛道未取得优于基线的结果。

原文摘要 · Abstract (English)

In this report, we describe the speaker diarization (SD) and language diarization (LD) systems developed by our team for the Second DISPLACE Challenge, 2024. Our contributions were dedicated to Track 1 for SD and Track 2 for LD in multilingual and multi-speaker scenarios. We investigated different speech enhancement techniques, voice activity detection (VAD) techniques, unsupervised domain categorization, and neural embedding extraction architectures. We also exploited the fusion of various embedding extraction models. We implemented our system with the open-source SpeechBrain toolkit. Our final submissions use spectral clustering for both the speaker and language diarization. We achieve about $7\%$ relative improvement over the challenge baseline in Track 1. We did not obtain improvement over the challenge baseline in Track 2.

语音分段语言分段多语种谱聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。