用状态空间模型统一处理多种语音失真,兼顾回归与生成两种方式提升泛化能力。
Universal Speech Enhancement with Regression and Generative Mamba
- 采用状态空间结构建模长序列,支持时频特征和采样率无关的表示学习。
- 在七类失真、五种语言下实现2nd名,盲测中表现优异。
- 针对丢包和带宽扩展场景设计生成式模块,弥补缺失内容更有效。
Interspeech 2025 URGENT 挑战旨在推动通用、鲁棒且可泛化的语音增强技术,涵盖七种不同失真类型和五种语言。我们提出通用语音增强马比(USEMamba),一种基于状态空间的语音增强模型,具备长序列建模、时频结构化处理及采样频率无关特征提取能力。方法以回归建模为主,在多数失真条件下表现良好;而对需推断缺失内容的丢包与带宽扩展场景,采用生成式变体更为有效。尽管仅使用部分训练数据,USEMamba在盲测阶段仍取得Track 1第2名,展现出强大跨条件泛化能力。
原文摘要 · Abstract (English)
The Interspeech 2025 URGENT Challenge aimed to advance universal, robust, and generalizable speech enhancement by unifying speech enhancement tasks across a wide variety of conditions, including seven different distortion types and five languages. We present Universal Speech Enhancement Mamba (USEMamba), a state-space speech enhancement model designed to handle long-range sequence modeling, time-frequency structured processing, and sampling frequency-independent feature extraction. Our approach primarily relies on regression-based modeling, which performs well across most distortions. However, for packet loss and bandwidth extension, where missing content must be inferred, a generative variant of the proposed USEMamba proves more effective. Despite being trained on only a subset of the full training data, USEMamba achieved 2nd place in Track 1 during the blind test phase, demonstrating strong generalization across diverse conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。