arXiv:2501.03184eess.AScs.LG2025-01被引 1

用自监督学习提升噪声下目标说话人语音活动检测精度

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining

  • 提出因果自监督框架DN-APC,利用无标签数据预训练
  • 在已知和未知噪声环境中性能提升约2%
  • FiLM条件控制方法表现最佳,适合嘈杂环境应用

目标说话人语音活动检测(TS-VAD)旨在识别音频帧中特定目标说话人的语音存在。尽管基于深度神经网络的模型表现良好,但其训练需大量标注数据,获取成本高,尤其在需泛化至未见环境时更显困难。为此,本文提出一种因果自监督学习(SSL)预训练框架——去噪自回归预测编码(DN-APC),以增强模型在噪声环境下的性能。同时,探索多种说话人条件化方法,并在不同噪声条件下评估其表现。实验表明,DN-APC在已知与未知噪声环境下均带来约2%的性能提升。此外,发现FiLM条件化方法整体表现最优。通过tSNE可视化分析,预训练后语音与非语音表示具有更强鲁棒性,验证了SSL预训练在提升TS-VAD模型抗噪能力方面的有效性。

原文摘要 · Abstract (English)

Target-Speaker Voice Activity Detection (TS-VAD) is the task of detecting the presence of speech from a known target-speaker in an audio frame. Recently, deep neural network-based models have shown good performance in this task. However, training these models requires extensive labelled data, which is costly and time-consuming to obtain, particularly if generalization to unseen environments is crucial. To mitigate this, we propose a causal, Self-Supervised Learning (SSL) pretraining framework, called Denoising Autoregressive Predictive Coding (DN-APC), to enhance TS-VAD performance in noisy conditions. We also explore various speaker conditioning methods and evaluate their performance under different noisy conditions. Our experiments show that DN-APC improves performance in noisy conditions, with a general improvement of approx. 2% in both seen and unseen noise. Additionally, we find that FiLM conditioning provides the best overall performance. Representation analysis via tSNE plots reveals robust initial representations of speech and non-speech from pretraining. This underscores the effectiveness of SSL pretraining in improving the robustness and performance of TS-VAD models in noisy environments.

语音检测自监督学习噪声鲁棒说话人条件

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。