用频谱-瞬态建模实现高质量语音超分辨率,实时高效。
STSR: High-Fidelity Speech Super-Resolution via Spectral-Transient Context Modeling
- 在MDCT域构建统一框架,融合全局频谱与局部瞬态信息。
- 在48kHz下实现一致的谐波重建,显著提升听觉保真度。
- 适合需要实时高保真语音恢复的应用场景。
语音超分辨率(SR)旨在从低分辨率输入重建高保真宽带语音,需兼顾全局谐波一致性与局部瞬态清晰度。尽管基于扩散模型的方法能提供出色保真度,但其计算开销过大难以实用;而高效的时域架构则缺乏显式频域表示,难以捕捉长程谱依赖关系并确保精确谐波对齐。本文提出STSR,一种在MDCT域构建的统一端到端框架,通过谱上下文注意力机制,利用分层窗口化自适应聚合非局部谱上下文,实现高达48kHz的稳定谐波重建。同时,采用稀疏感知正则化策略缓解压缩谱表示中瞬态成分被抑制的问题。STSR在感知保真度和零样本泛化能力上持续优于现有最优基线,为高质量语音恢复提供了鲁棒、实时的新范式。
原文摘要 · Abstract (English)
Speech super-resolution (SR) reconstructs high-fidelity wideband speech from low-resolution inputs-a task that necessitates reconciling global harmonic coherence with local transient sharpness. While diffusion-based generative models yield impressive fidelity, their practical deployment is often stymied by prohibitive computational demands. Conversely, efficient time-domain architectures lack the explicit frequency representations essential for capturing long-range spectral dependencies and ensuring precise harmonic alignment. We introduce STSR, a unified end-to-end framework formulated in the MDCT domain to circumvent these limitations. STSR employs a Spectral-Contextual Attention mechanism that harnesses hierarchical windowing to adaptively aggregate non-local spectral context, enabling consistent harmonic reconstruction up to 48 kHz. Concurrently, a sparse-aware regularization strategy is employed to mitigate the suppression of transient components inherent in compressed spectral representations. STSR consistently outperforms state-of-the-art baselines in both perceptual fidelity and zero-shot generalization, providing a robust, real-time paradigm for high-quality speech restoration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。