arXiv:2604.01832eess.AScs.AI2026-04中稿 · ICASSP 2026被引 1

融合生成与预测的语音增强框架,提升鲁棒性与听感质量。

GAP-URGENet: A Generative-Predictive Fusion Framework for Universal Speech Enhancement

  • 分两路处理:生成分支在自监督表示域修复语音,预测分支在频谱域增强信号。
  • 盲测表现最优,客观评价排名第一,48kHz宽带输出后下采样至原采样率。
  • 适合语音增强、降噪场景,尤其在复杂噪声环境下表现突出。

我们提出GAP-URGENet,一种为ICASSP 2026 URGENT挑战赛第1赛道设计的生成-预测融合框架。系统包含生成分支,在自监督表示域完成全栈语音修复,并通过神经声码器重建波形;以及预测分支,在频谱域执行增强,提供互补信息。两分支输出由后处理模块融合,该模块还进行带宽扩展,生成48 kHz增强波形,随后下采样至原始采样率。该融合策略显著提升鲁棒性与听觉质量,在盲测阶段表现最优,客观评价排名第一。音频示例可访问 https://xiaobin-rong.github.io/gap-urgenet_demo。

原文摘要 · Abstract (English)

We introduce GAP-URGENet, a generative-predictive fusion framework developed for Track 1 of the ICASSP 2026 URGENT Challenge. The system integrates a generative branch, which performs full-stack speech restoration in a self-supervised representation domain and reconstructs the waveform via a neural vocoder, along with a predictive branch that performs spectrogram-domain enhancement, providing complementary cues. Outputs from both branches are fused by a post-processing module, which also performs bandwidth extension to generate the enhanced waveform at 48 kHz, later downsampled to the original sampling rate. This generative-predictive fusion improves robustness and perceptual quality, achieving top performance in the blind-test phase and ranking 1st in the objective evaluation. Audio examples are available at https://xiaobin-rong.github.io/gap-urgenet_demo.

语音增强生成模型融合框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。