用环形混合与新损失函数,让语音分离模型从嘈杂录音中自动去噪。
Ring Mixing with Auxiliary Signal-to-Consistency-Error Ratio Loss for Unsupervised Denoising in Speech Separation
- 设计环形混合策略,同一语音源参与两次混合,增强训练一致性
- 引入SCER辅助损失,使相同语音在不同混合中的估计更一致,降低残余噪声超50%
- 首次实现仅用真实嘈杂语音训练,提升模型对真实场景的泛化能力
现有噪声语音分离系统通常在全合成混合数据上训练,限制了在真实场景下的泛化能力。尽管可使用领域内(常含噪声)语音混合进行训练,但研究表明这会导致不良最优解:混合噪声被保留在输出估计中,原因在于背景噪声不可分且损失函数具有对称性。为此,本文提出环形混合策略,即每个语音源在两个不同混合中出现,并引入信号-一致性误差比(SCER)辅助损失,惩罚同一源在不同混合中不一致的估计,打破对称性,激励模型去噪。在基于WHAM!的基准测试中,该方法可使残余噪声减少超过一半,证明仅使用嘈杂录音即可有效学习去噪。该方法为利用真实世界数据训练更具泛化性的语音分离系统开辟了新路径,我们通过使用VoxCeleb中的自然嘈杂语音训练系统进行了验证。
原文摘要 · Abstract (English)
Noisy speech separation systems are typically trained on fully-synthetic mixtures, limiting generalization to real-world scenarios. Though training on mixtures of in-domain (thus often noisy) speech is possible, we show that this leads to undesirable optima where mixture noise is retained in the estimates, due to the inseparability of the background noises and the loss function's symmetry. To address this, we propose ring mixing, a batch strategy of using each source in two mixtures, alongside a new Signal-to-Consistency-Error Ratio (SCER) auxiliary loss penalizing inconsistent estimates of the same source from different mixtures, breaking symmetry and incentivizing denoising. On a WHAM!-based benchmark, our method can reduce residual noise by upwards of half, effectively learning to denoise from only noisy recordings. This opens the door to training more generalizable systems using in-the-wild data, which we demonstrate via systems trained using naturally-noisy speech from VoxCeleb.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。