arXiv:2506.00809cs.SDeess.AS2025-06中稿 · INTERSPEECH 2025被引 3

用多阶段融合提升语音增强效果,适配多种噪声和语言。

FUSE: Universal Speech Enhancement using Multi-Stage Fusion of Sparse Compression and Token Generation Models for the URGENT 2025 Challenge

  • 先用稀疏压缩分离噪声,再用生成模型优化语音质量。
  • 在多语言数据上提升信号保真度与听感,表现优于基线。
  • 适合需要高鲁棒性语音增强的工业场景应用。

我们提出一种多阶段通用语音增强框架,专为2025年Interspeech URGENT挑战设计。系统首先通过稀疏压缩网络从含噪输入中鲁棒地分离声源并提取初始纯净语音估计;随后,利用自监督特征和神经音频编解码器生成的声学标记,基于掩码语言建模目标优化生成模型,进一步提升语音质量;最后,融合网络将前两阶段输出与原始含噪信号结合,实现信号保真度与感知质量的平衡提升。此外,采用时间偏移聚合技巧及输出融合策略,进一步增强性能。在具有不同采样率和多样失真类型的多语言挑战数据集上的实验验证了方法的有效性。

原文摘要 · Abstract (English)

We propose a multi-stage framework for universal speech enhancement, designed for the Interspeech 2025 URGENT Challenge. Our system first employs a Sparse Compression Network to robustly separate sources and extract an initial clean speech estimate from noisy inputs. This is followed by an efficient generative model that refines speech quality by leveraging self-supervised features and optimizing a masked language modeling objective on acoustic tokens derived from a neural audio codec. In the final stage, a fusion network integrates the outputs of the first two stages with the original noisy signal, achieving a balanced improvement in both signal fidelity and perceptual quality. Additionally, a shift trick that aggregates multiple time-shifted predictions, along with output blending, further boosts performance. Experimental results on challenging multilingual datasets with variable sampling rates and diverse distortion types validate the effectiveness of our approach.

语音增强多阶段融合自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。