提出复数域注意力网络,同时重建语音幅度与相位,提升超分辨率质量。
A High-Fidelity Speech Super Resolution Network using a Complex Global Attention Module with Spectro-Temporal Loss
- 在复数域建模幅度与相位,引入复杂全局注意力捕捉依赖关系
- 在VCTK数据集上2kHz→48kHz超采样下优于主流模型,无噪声伪影
- 适合语音增强、低码率语音修复等需要高保真输出的场景
语音超分辨率(SSR)通过提高采样率增强低分辨率语音。现有方法多关注幅度重建,但近期研究指出相位重建对感知质量至关重要。为此,本文提出CTFT-Net:一种在复数域同时重建幅度与相位的复杂时频变换网络。该网络采用复数全局注意力模块建模音素间与频段间的依赖关系,并结合复数Conformer捕获长程与局部特征,提升频率重建精度与抗噪能力。此外,设计了时域与多分辨率频域损失函数以增强泛化性。实验表明,CTFT-Net在VCTK数据集上显著优于当前最优模型(NU-Wave、WSRGlow、NVSR、AERO),尤其在极端超采样(2 kHz → 48 kHz)条件下,有效恢复高频成分且无噪声伪影。
原文摘要 · Abstract (English)
Speech super-resolution (SSR) enhances low-resolution speech by increasing the sampling rate. While most SSR methods focus on magnitude reconstruction, recent research highlights the importance of phase reconstruction for improved perceptual quality. Therefore, we introduce CTFT-Net, a Complex Time-Frequency Transformation Network that reconstructs both magnitude and phase in complex domains for improved SSR tasks. It incorporates a complex global attention block to model inter-phoneme and inter-frequency dependencies and a complex conformer to capture long-range and local features, improving frequency reconstruction and noise robustness. CTFT-Net employs time-domain and multi-resolution frequency-domain loss functions for better generalization. Experiments show CTFT-Net outperforms state-of-the-art models (NU-Wave, WSRGlow, NVSR, AERO) on the VCTK dataset, particularly for extreme upsampling (2 kHz to 48 kHz), reconstructing high frequencies effectively without noisy artifacts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。