直接在时域完成语音超分辨率,速度更快、质量更高。
Wave-U-Mamba: An End-To-End Framework For High-Quality And Efficient Speech Super Resolution
- 直接在时域重建语音,避免频谱相位信息丢失
- 在8-24kHz低采样率下均达最低频谱距离,主观评分最优
- 推理速度超基线9倍,参数量小于1/50,适合部署
语音超分辨率(SSR)旨在通过恢复缺失的高频成分来提升低分辨率语音信号。传统方法通常先重建对数梅尔频谱,再通过声码器在波形域生成高分辨率语音,但梅尔特征缺乏相位信息,可能导致重建性能下降。受选择性状态空间模型(SSMs)启发,我们提出直接在时域进行SSR的Wave-U-Mamba框架。对比实验显示,该方法在多种低采样率(8–24 kHz)下均取得最低的对数谱距离(LSD),显著优于WSRGlow、NU-Wave 2和AudioSR等基线模型。主观评测(使用平均意见分,MOS)表明,生成语音自然度高、类人。此外,其单张A100 GPU上的推理速度比基线快九倍以上,参数量不足基线模型的2%。
原文摘要 · Abstract (English)
Speech Super-Resolution (SSR) is a task of enhancing low-resolution speech signals by restoring missing high-frequency components. Conventional approaches typically reconstruct log-mel features, followed by a vocoder that generates high-resolution speech in the waveform domain. However, as mel features lack phase information, this can result in performance degradation during the reconstruction phase. Motivated by recent advances with Selective State Spaces Models (SSMs), we propose a method, referred to as Wave-U-Mamba that directly performs SSR in time domain. In our comparative study, including models such as WSRGlow, NU-Wave 2, and AudioSR, Wave-U-Mamba demonstrates superior performance, achieving the lowest Log-Spectral Distance (LSD) across various low-resolution sampling rates, ranging from 8 to 24 kHz. Additionally, subjective human evaluations, scored using Mean Opinion Score (MOS) reveal that our method produces SSR with natural and human-like quality. Furthermore, Wave-U-Mamba achieves these results while generating high-resolution speech over nine times faster than baseline models on a single A100 GPU, with parameter sizes less than 2\% of those in the baseline models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。