arXiv:2501.07897cs.SDeess.AS2025-01中稿 · ICASSP 2025被引 15

用薛定谔桥实现高效语音超分辨率,1.7万参数模型更快更清。

Bridge-SR: Schrödinger Bridge for Efficient SR

  • 基于薛定谔桥建模低频到高频波形的生成过程。
  • 4步合成比8步扩散模型更快,音质误差(LSD)更低(0.911 vs 0.927)。
  • 轻量网络+优化噪声调度,适合实时语音修复场景。

语音超分辨率(SR)旨在从低采样率波形重建高采样率版本,是语音恢复的关键任务。现有方法多在不同数据空间中尝试,但常需额外压缩网络或质量与速度受限。本文提出Bridge-SR,一种在波形域中高效的任意输入至48kHz的语音超分辨率系统。利用可解析的薛定谔桥模型,将观测到的低分辨率波形作为先验,其本身蕴含目标高分辨率信号的丰富信息。通过优化轻量网络学习从先验到目标的得分函数,实现端到端的数据生成,充分挖掘低频观测中的指导内容。进一步发现噪声调度、数据缩放及辅助损失函数对性能至关重要。在基准数据集VCTK上的实验表明:(1)使用仅170万参数的轻量骨干网络,各项指标优于多个强基线;(2)4步合成在音质(LSD: 0.911)和推理速度上均优于8步条件扩散模型(LSD: 0.927)。演示见https://bridge-sr.github.io。

原文摘要 · Abstract (English)

Speech super-resolution (SR), which generates a waveform at a higher sampling rate from its low-resolution version, is a long-standing critical task in speech restoration. Previous works have explored speech SR in different data spaces, but these methods either require additional compression networks or exhibit limited synthesis quality and inference speed. Motivated by recent advances in probabilistic generative models, we present Bridge-SR, a novel and efficient any-to-48kHz SR system in the speech waveform domain. Using tractable Schrödinger Bridge models, we leverage the observed low-resolution waveform as a prior, which is intrinsically informative for the high-resolution target. By optimizing a lightweight network to learn the score functions from the prior to the target, we achieve efficient waveform SR through a data-to-data generation process that fully exploits the instructive content contained in the low-resolution observation. Furthermore, we identify the importance of the noise schedule, data scaling, and auxiliary loss functions, which further improve the SR quality of bridge-based systems. The experiments conducted on the benchmark dataset VCTK demonstrate the efficiency of our system: (1) in terms of sample quality, Bridge-SR outperforms several strong baseline methods under different SR settings, using a lightweight network backbone (1.7M); (2) in terms of inference speed, our 4-step synthesis achieves better performance than the 8-step conditional diffusion counterpart (LSD: 0.911 vs 0.927). Demo at https://bridge-sr.github.io.

语音超分辨薛定谔桥轻量化波形生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。