arXiv:2512.09713eess.AS2025-12

提出抗歌声干扰的语音活动检测模型,提升复杂场景下语音识别准确性

Robust Speech Activity Detection in the Presence of Singing Voice

  • 通过控制语音与歌声样本比例训练模型,增强两者区分能力
  • 在多音乐风格数据集上实现0.919的AUC,有效拒绝不属于语音的歌声
  • 适合语音识别、对话增强等需精准识别语音的场景

语音活动检测(SAD)系统常将歌声误判为语音,影响对话增强和自动语音识别性能。本文提出歌唱鲁棒语音活动检测(SR-SAD),一种神经网络模型,可在存在歌声时仍准确检测语音。主要贡献包括:(i) 使用可控比例的语音与歌声样本进行训练,提升二者的区分能力;(ii) 设计计算高效的模型,在保持鲁棒性的同时降低推理耗时;(iii) 提出新评估指标,专门用于衡量混合语音-歌声场景下的SAD鲁棒性。在涵盖多种音乐风格的挑战性数据集上的实验表明,SR-SAD在保持高语音检测精度(AUC = 0.919)的同时能有效排除歌声。通过显式学习区分语音与歌声,该模型显著提升了混合场景中的语音活动检测可靠性。

原文摘要 · Abstract (English)

Speech Activity Detection (SAD) systems often misclassify singing as speech, leading to degraded performance in applications such as dialogue enhancement and automatic speech recognition. We introduce Singing-Robust Speech Activity Detection ( SR-SAD ), a neural network designed to robustly detect speech in the presence of singing. Our key contributions are: i) a training strategy using controlled ratios of speech and singing samples to improve discrimination, ii) a computationally efficient model that maintains robust performance while reducing inference runtime, and iii) a new evaluation metric tailored to assess SAD robustness in mixed speech-singing scenarios. Experiments on a challenging dataset spanning multiple musical genres show that SR-SAD maintains high speech detection accuracy (AUC = 0.919) while rejecting singing. By explicitly learning to distinguish between speech and singing, SR-SAD enables more reliable SAD in mixed speech-singing scenarios.

语音检测歌声干扰深度学习语音识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。