统一生成式语音增强与分离,提升复杂降噪下的语音质量。
Geneses: Unified Generative Speech Enhancement and Separation
- 用多模态扩散Transformer和潜在流匹配估计清晰语音特征
- 在两种噪声条件下均显著优于传统掩码方法
- 适合需要高鲁棒性语音处理的场景
真实世界音频常包含多个说话人及多种退化问题,限制了高质量语音数据的可用性。尽管端到端的语音增强(SE)与分离(SS)联合方法具有潜力,但传统方法对加性噪声以外的复杂退化表现不佳。为此,我们提出 extbf{Geneses},一种统一的生成式框架,实现高质量的语音增强与分离。Geneses 利用潜在流匹配,基于自监督学习表示从含噪混合信号中估计每个说话人的清晰语音特征。我们在 LibriTTS-R 的双说话人混合数据上进行评估,分别在仅加性噪声和复杂退化两种条件下测试。结果表明,Geneses 在多项客观指标上显著优于传统的基于掩码的 SE--SS 方法,且对复杂退化具有强鲁棒性。音频样例可在演示页面获取。
原文摘要 · Abstract (English)
Real-world audio recordings often contain multiple speakers and various degradations, which limit both the quantity and quality of speech data available for building state-of-the-art speech processing models. Although end-to-end approaches that concatenate speech enhancement (SE) and speech separation (SS) to obtain a clean speech signal for each speaker are promising, conventional SE-SS methods suffer from complex degradations beyond additive noise. To this end, we propose \textbf{Geneses}, a generative framework to achieve unified, high-quality SE--SS. Our Geneses leverages latent flow matching to estimate each speaker's clean speech features using multi-modal diffusion Transformer conditioned on self-supervised learning representation from noisy mixture. We conduct experimental evaluation using two-speaker mixtures from LibriTTS-R under two conditions: additive-noise-only and complex degradations. The results demonstrate that Geneses significantly outperforms a conventional mask-based SE--SS method across various objective metrics with high robustness against complex degradations. Audio samples are available in our demo page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。