三阶段设计提升语音增强对多种失真的鲁棒性与泛化能力
TS-URGENet: A Three-stage Universal Robust and Generalizable Speech Enhancement Network
- 分填空、分离、修复三阶段处理不同语音失真
- 在Interspeech 2025挑战赛中排名第二
- 适合需要跨场景通用语音增强的系统开发者
通用语音增强旨在处理具有不同失真和输入格式的语音信号。为此,我们提出TS-URGENet——一种三阶段通用、鲁棒且可泛化的语音增强网络。该系统采用新颖的三阶段架构:填空阶段在噪声干扰下初步填补丢失区域,保障信号连续性;分离阶段抑制噪声、混响和削波失真,提升语音清晰度;修复阶段补偿带宽限制、编解码器伪影及残留包丢失失真,优化整体语音质量。所提方法在Interspeech 2025 URGENT挑战赛中表现优异,Track 1排名第二。
原文摘要 · Abstract (English)
Universal speech enhancement aims to handle input speech with different distortions and input formats. To tackle this challenge, we present TS-URGENet, a Three-Stage Universal, Robust, and Generalizable speech Enhancement Network. To address various distortions, the proposed system employs a novel three-stage architecture consisting of a filling stage, a separation stage, and a restoration stage. The filling stage mitigates packet loss by preliminarily filling lost regions under noise interference, ensuring signal continuity. The separation stage suppresses noise, reverberation, and clipping distortion to improve speech clarity. Finally, the restoration stage compensates for bandwidth limitation, codec artifacts, and residual packet loss distortion, refining the overall speech quality. Our proposed TS-URGENet achieved outstanding performance in the Interspeech 2025 URGENT Challenge, ranking 2nd in Track 1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。