arXiv:2505.23207cs.SDcs.LG2025-05被引 2

用WavLM和说话人注意力提升重叠语音检测精度

Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM

  • 分阶段训练增强语音活动与重叠检测的关联性
  • AMI数据集上达到82.76%的F1分数,性能领先
  • 适合做多说话人语音处理的工程师和研究者

重叠语音检测(OSD)旨在识别对话中多个说话人重叠的区域,是多人语音处理的关键挑战。本文提出一种基于说话人感知的渐进式OSD模型,采用渐进式训练策略提升语音活动检测(VAD)与重叠检测之间的相关性。为优化声学表征,探索了先进自监督学习(SSL)模型WavLM和wav2vec 2.0的性能,并引入帧级说话人注意力模块以增强特征。实验结果表明,该方法在AMI测试集上取得了82.76%的F1分数,验证了其在OSD任务中的鲁棒性与有效性。

原文摘要 · Abstract (English)

Overlapping Speech Detection (OSD) aims to identify regions where multiple speakers overlap in a conversation, a critical challenge in multi-party speech processing. This work proposes a speaker-aware progressive OSD model that leverages a progressive training strategy to enhance the correlation between subtasks such as voice activity detection (VAD) and overlap detection. To improve acoustic representation, we explore the effectiveness of state-of-the-art self-supervised learning (SSL) models, including WavLM and wav2vec 2.0, while incorporating a speaker attention module to enrich features with frame-level speaker information. Experimental results show that the proposed method achieves state-of-the-art performance, with an F1 score of 82.76\% on the AMI test set, demonstrating its robustness and effectiveness in OSD.

语音检测重叠语音自监督学习说话人信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。