arXiv:2501.11311cs.SDcs.LG2025-01被引 11

A2SB可端到端修复高保真音乐,支持频段扩展与缺失段重建。

A2SB: Audio-to-Audio Schrodinger Bridges

  • 基于薛定谔桥框架,直接生成波形无需声码器
  • 在44.1kHz下实现小时级音频的频段扩展与补全
  • 适用于无版权限制音乐数据,适合高保真音频修复场景

真实世界音频常受多种因素影响而退化。本文提出针对44.1kHz高保真音乐的音频修复模型A2SB,能够同时实现带宽扩展(预测高频成分)和插补(重建缺失片段)。关键优势在于端到端设计,无需声码器即可输出波形,支持小时级音频输入,并使用宽松许可音乐数据训练。A2SB在多个分布外音乐测试集上实现了最先进的带宽扩展与插补质量。

原文摘要 · Abstract (English)

Real-world audio is often degraded by numerous factors. This work presents an audio restoration model tailored for high-res music at 44.1kHz. Our model, Audio-to-Audio Schrödinger Bridges (A2SB), is capable of both bandwidth extension (predicting high-frequency components) and inpainting (re-generating missing segments). Critically, A2SB is end-to-end requiring no vocoder to predict waveform outputs, able to restore hour-long audio inputs, and trained on permissively licensed music data. A2SB is capable of achieving state-of-the-art band-width extension and inpainting quality on several out-of-distribution music test sets.

音频修复波形生成薛定谔桥高保真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。