无需配对数据,用无监督方法分离语音与噪声并提升音质
Unsupervised Speech Enhancement using Data-defined Priors
- 双分支结构分别建模干净语音和残留噪声
- 利用对抗训练引入未配对数据定义的先验,实现无监督优化
- 发现使用同领域数据定义先验会夸大性能,需谨慎选择数据
大多数基于深度学习的语音增强方法依赖成对的清晰-噪声语音数据。在真实场景下大规模收集此类数据不现实,因此社区普遍采用合成噪声语音。但这造成了训练与测试阶段的分布差异。本文提出一种新颖的无监督双分支编码器-解码器架构,将输入分离为干净语音和残余噪声。通过对抗训练,在未配对的清晰语音和可选噪声数据上施加先验。实验表明,该方法性能可媲美当前领先的无监督语音增强技术。此外,我们揭示了清晰语音数据选择对增强效果的关键影响:若在域内数据上定义先验,性能可能被过度乐观估计——这是以往无监督研究中常见的做法。
原文摘要 · Abstract (English)
The majority of deep learning-based speech enhancement methods require paired clean-noisy speech data. Collecting such data at scale in real-world conditions is infeasible, which has led the community to rely on synthetically generated noisy speech. However, this introduces a gap between the training and testing phases. In this work, we propose a novel dual-branch encoder-decoder architecture for unsupervised speech enhancement that separates the input into clean speech and residual noise. Adversarial training is employed to impose priors on each branch, defined by unpaired datasets of clean speech and, optionally, noise. Experimental results show that our method achieves performance comparable to leading unsupervised speech enhancement approaches. Furthermore, we demonstrate the critical impact of clean speech data selection on enhancement performance. In particular, our findings reveal that performance may appear overly optimistic when in-domain clean speech data are used for prior definition -- a practice adopted in previous unsupervised speech enhancement studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。