用频谱渐进学习提升语音宽带重建的音质与泛化能力
FSC-Net: Integrating Fast Fourier Convolutions and Progressive Learning for Speech Bandwidth Extension

- 引入快速傅里叶卷积捕捉全频段长程依赖
- 在VCTK数据集上4kHz到48kHz任务中达最优音质指标
- 参数量仅154万,适合资源受限场景部署
语音宽带扩展(BWE)旨在从窄带输入重建高保真宽带音频。尽管近期方法取得进展,但常难以重建真实高频相位与谐波结构,导致听觉失真。本文提出FSC-Net(全频谱上下文网络),一种参数高效架构,显式建模跨频带谐波依赖。通过将快速傅里叶卷积(FFCs)融入复数谱映射框架,FSC-Net将感受野扩展至整个频谱,有效捕捉长程频率交互。为应对高频生成的病态性,提出新颖的频率渐进学习课程,引导网络由粗到细重建谱细节。在VCTK和未见的EARS数据集上的实验表明,FSC-Net在挑战性的4 kHz至48 kHz任务中持续表现优异,相较放大基线模型,以更小参数量(1.54 M)获得领先的小波谱失真(LSD)与客观音质评分(PESQ)。
原文摘要 · Abstract (English)
Speech bandwidth extension (BWE) aims to reconstruct high-fidelity wideband audio from narrowband inputs. While recent approaches have made significant progress, they often struggle to reconstruct realistic high-frequency phase and harmonic structures, leading to perceptual artifacts. In this paper, we propose FSC-Net (Full-Spectrum Context Network), a parameter-efficient architecture designed to explicitly model cross-band harmonic dependencies. By integrating Fast Fourier Convolutions (FFCs) into a complex spectral mapping framework, FSC-Net expands its receptive field to the entire spectrum, capturing long-range frequency interactions effectively. To address the ill-posed nature of high-frequency generation, our novel frequency-progressive learning curriculum guides the network to reconstruct spectral details from coarse to fine. Experimental results on the VCTK and unseen EARS datasets demonstrate that FSC-Net delivers consistently strong reconstruction quality and generalization, particularly in the challenging VCTK 4 kHz-to-48 kHz task. Compared to scaled-up baselines, our model attains leading LSD and PESQ scores while maintaining a highly compact parameter footprint (1.54 M).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。