arXiv:2508.20885cs.SD2025-08中稿 · IEEE ASRU 2025被引 3

用可学习滤波器和排序优化提升语音活动检测在噪声下的鲁棒性。

SincQDR-VAD: A Noise-Robust Voice Activity Detection Framework Leveraging Learnable Filters and Ranking-Aware Optimization

  • 采用可学习带通滤波器提取抗噪频谱特征
  • 引入二次差异排序损失,使语音/非语音帧得分排序更优,AUROC显著提升
  • 参数量仅需前人方法的69%,适合资源受限场景

语音活动检测(VAD)对语音驱动应用至关重要,但在噪声和资源受限环境下仍不理想。现有方法通常对噪声敏感,且帧级分类损失与评估指标关联松散。为此,我们提出SincQDR-VAD,一种紧凑且鲁棒的框架,结合Sinc提取器前端与新型二次差异排序损失。Sinc提取器使用可学习带通滤波器捕捉抗噪频谱特征,而排序损失通过优化语音与非语音帧间的得分顺序,提升受试者工作特征曲线下面积(AUROC)。在代表性基准数据集上的实验表明,该框架显著提升AUROC与F2-Score,同时仅需前人方法69%的参数量,验证了其高效性与实用性。

原文摘要 · Abstract (English)

Voice activity detection (VAD) is essential for speech-driven applications, but remains far from perfect in noisy and resource-limited environments. Existing methods often lack robustness to noise, and their frame-wise classification losses are only loosely coupled with the evaluation metric of VAD. To address these challenges, we propose SincQDR-VAD, a compact and robust framework that combines a Sinc-extractor front-end with a novel quadratic disparity ranking loss. The Sinc-extractor uses learnable bandpass filters to capture noise-resistant spectral features, while the ranking loss optimizes the pairwise score order between speech and non-speech frames to improve the area under the receiver operating characteristic curve (AUROC). A series of experiments conducted on representative benchmark datasets show that our framework considerably improves both AUROC and F2-Score, while using only 69% of the parameters compared to prior arts, confirming its efficiency and practical viability.

语音检测抗噪模型排序优化轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。