arXiv:2508.14623eess.AScs.AI2025-08中稿 · IEEE ASRU 2025, Wo…被引 5

改进语音分离中含噪参考信号的评估与训练,降低输出噪声。

A Study of the Scale Invariant Signal to Distortion Ratio in Speech Separation with Noisy References

  • 用增强参考信号和WHAM!数据增强混合音频,避免模型学习噪声。
  • 实验显示分离语音噪声减少,但处理参考信号可能引入伪影。
  • 发现SI-SDR与听感噪声负相关,验证方法有效性。

本文研究在训练参考信号含噪声(如标准数据集WSJ0-2Mix)时,使用尺度不变信干比(SI-SDR)作为评价与训练目标的影响。推导表明,噪声会限制可达到的SI-SDR值,或导致分离输出中出现不期望的噪声。为此,提出一种增强参考信号并结合WHAM!数据增强混合信号的方法,旨在训练出避免学习噪声的模型。在两个模型上使用非侵入式指标NISQA.v2进行评估,结果表明分离语音噪声降低,但参考信号处理可能引入伪影,限制整体质量提升。在WSJ0-2Mix和Libri2Mix测试集上,所有模型均呈现SI-SDR与感知噪声间的负相关,验证了理论推导结论。

原文摘要 · Abstract (English)

This paper examines the implications of using the Scale-Invariant Signal-to-Distortion Ratio (SI-SDR) as both evaluation and training objective in supervised speech separation, when the training references contain noise, as is the case with the de facto benchmark WSJ0-2Mix. A derivation of the SI-SDR with noisy references reveals that noise limits the achievable SI-SDR, or leads to undesired noise in the separated outputs. To address this, a method is proposed to enhance references and augment the mixtures with WHAM!, aiming to train models that avoid learning noisy references. Two models trained on these enhanced datasets are evaluated with the non-intrusive NISQA.v2 metric. Results show reduced noise in separated speech but suggest that processing references may introduce artefacts, limiting overall quality gains. Negative correlation is found between SI-SDR and perceived noisiness across models on the WSJ0-2Mix and Libri2Mix test sets, underlining the conclusion from the derivation.

语音分离评估指标噪声处理训练优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。