arXiv:2503.12936eess.AS2025-03被引 6

用真实混响语音训练语音增强模型,显著降低错误率。

FNSE-SBGAN: Far-field Speech Enhancement with Schrodinger Bridge and Generative Adversarial Networks

  • 结合薛定谔桥与生成对抗网络,直接在真实数据上训练
  • 相比远场信号,字符错误率最高降低14.58%
  • 适合需要高保真语音增强的现实场景应用

现有神经语音增强方法多依赖模拟的远场噪声混响语音与干净语音配对数据,但泛化能力受限于真实场景。本文研究直接在真实混合语音上训练增强模型,聚焦低信噪比、强混响及中高频衰减的真实环境下的单通道远场到近场语音增强(FNSE)任务。提出FNSE-SBGAN框架,融合基于薛定谔桥(SB)的扩散模型与生成对抗网络(GAN)。实验表明,该方法在多项指标和主观评价中均达到当前最优,字符错误率(CER)最高降低14.58%。此外,引入基于时频域矩阵秩分析的评估框架,系统揭示不同生成方法的优劣,为模型性能提供深入洞察。

原文摘要 · Abstract (English)

The prevailing method for neural speech enhancement predominantly utilizes fully-supervised deep learning with simulated pairs of far-field noisy-reverberant speech and clean speech. Nonetheless, these models frequently demonstrate restricted generalizability to mixtures recorded in real-world conditions. To address this issue, this study investigates training enhancement models directly on real mixtures. Specifically, we revisit the single-channel far-field to near-field speech enhancement (FNSE) task, focusing on real-world data characterized by low signal-to-noise ratio (SNR), high reverberation, and mid-to-high frequency attenuation. We propose FNSE-SBGAN, a framework that integrates a Schrodinger Bridge (SB)-based diffusion model with generative adversarial networks (GANs). Our approach achieves state-of-the-art performance across various metrics and subjective evaluations, significantly reducing the character error rate (CER) by up to 14.58% compared to far-field signals. Experimental results demonstrate that FNSE-SBGAN preserves superior subjective quality and establishes a new benchmark for real-world far-field speech enhancement. Additionally, we introduce an evaluation framework leveraging matrix rank analysis in the time-frequency domain, providing systematic insights into model performance and revealing the strengths and weaknesses of different generative methods.

语音增强扩散模型真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。