arXiv:2507.08333cs.SDcs.AI2025-07被引 2

用离散扩散模型修复音乐音频长段缺失,效果更稳定连贯。

Token-Based Audio Inpainting via Discrete Diffusion

  • 在预训练音频分词器的离散表示上应用扩散模型
  • 750毫秒内缺失段修复效果优于现有方法
  • 适合需要高质量音乐修复的研究与应用

音频修复旨在恢复退化录音中的缺失片段。以往基于扩散的方法在缺失区域较大时性能下降。本文提出首个在预训练音频分词器生成的离散音乐表示上应用离散扩散的方法,实现了长间隙的稳定且语义连贯的修复。方法引入两种训练策略:基于导数的正则化损失以保证时间动态平滑,以及基于区间吸收转移机制,在扩散过程中提供结构化破坏。在MusicNet和MAESTRO数据集上,针对长达750毫秒的缺失段进行实验,结果表明本方法在150毫秒及以上各类缺失长度下均持续优于强基线。该工作推动了音乐音频修复的发展,并为离散扩散模型训练提供了新方向。项目页面提供示例与代码。

原文摘要 · Abstract (English)

Audio inpainting seeks to restore missing segments in degraded recordings. Previous diffusion-based methods exhibit impaired performance when the missing region is large. We introduce the first approach that applies discrete diffusion over tokenized music representations from a pre-trained audio tokenizer, enabling stable and semantically coherent restoration of long gaps. Our method further incorporates two training approaches: a derivative-based regularization loss that enforces smooth temporal dynamics, and a span-based absorbing transition that provides structured corruption during diffusion. Experiments on the MusicNet and MAESTRO datasets with gaps up to 750 ms show that our approach consistently outperforms strong baselines across range of gap lengths, for gaps of 150 ms and above. This work advances musical audio restoration and introduces new directions for discrete diffusion model training. Visit our project page for examples and code.

音频修复扩散模型离散生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。