融合判别与生成模型,提升多场景语音增强效果
A Hybrid Discriminative and Generative System for Universal Speech Enhancement
- 结合判别与生成模型,兼顾信号保真与细节重建
- 支持不同采样率输入,实现通用语音增强
- 在ICASSP 2026挑战赛中位列第三,性能优异
通用语音增强旨在处理各种语音失真和录音条件。本文提出一种新型混合架构,融合判别建模的信号保真性与生成建模的重建能力。系统采用具有采样频率无关策略的判别型TF-GridNet模型,可统一处理不同采样率输入;同时,结合自回归模型与频谱映射的生成模块,生成高细节语音并有效抑制生成伪影。最后,融合网络在信号级损失与综合语音质量评估(SQA)损失的优化下,学习两路输出的自适应权重。该系统在ICASSP 2026 URGENT Challenge(Track 1)中排名第三。
原文摘要 · Abstract (English)
Universal speech enhancement aims at handling inputs with various speech distortions and recording conditions. In this work, we propose a novel hybrid architecture that synergizes the signal fidelity of discriminative modeling with the reconstruction capabilities of generative modeling. Our system utilizes the discriminative TF-GridNet model with the Sampling-Frequency-Independent strategy to handle variable sampling rates universally. In parallel, an autoregressive model combined with spectral mapping modeling generates detail-rich speech while effectively suppressing generative artifacts. Finally, a fusion network learns adaptive weights of the two outputs under the optimization of signal-level losses and the comprehensive Speech Quality Assessment (SQA) loss. Our proposed system is evaluated in the ICASSP 2026 URGENT Challenge (Track 1) and ranks the third place.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。