提出新框架提升语音去噪,兼顾全局与周期性频谱特征。
SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns
- 设计GLP模块,融合频带的全局、局部和周期特性。
- 采用多分辨率时频并行处理,捕捉多样频谱模式。
- 在保持高效的同时显著提升语音还原质量。
通用语音恢复需要能解析复杂语音结构的技术,尤其在多种失真条件下。尽管状态空间模型如SEMamba已在语音降噪中达到先进水平,但其未针对语音的关键特性(如频谱周期性或多分辨率频率分析)进行优化。本文提出一种结合语音特性的架构,引入全局、局部与周期(GLP)模块,作为频率特征提取单元,有效利用频带属性。同时设计多分辨率并行时频双处理块以捕获多样化频谱模式,并加入可学习映射进一步提升性能。实验表明,结合所有创新后,SEMamba++在多个基线模型中表现最佳,且计算效率高。
原文摘要 · Abstract (English)
General speech restoration demands techniques that can interpret complex speech structures under various distortions. While State-Space Models like SEMamba have advanced the state-of-the-art in speech denoising, they are not inherently optimized for critical speech characteristics, such as spectral periodicity or multi-resolution frequency analysis. In this work, we introduce an architecture tailored to incorporate speech-specific features as inductive biases. In particular, we propose the Global, Local, and Periodic (GLP) module, a frequency feature extraction block that effectively and efficiently leverages the properties of frequency bins. Then, we design a multi-resolution parallel time-frequency dual-processing block to capture diverse spectral patterns, and a learnable mapping to further enhance model performance. With all our ideas combined, the proposed SEMamba++ achieves the best performance among multiple baseline models while remaining computationally efficient.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。