arXiv:2509.16342eess.AScs.LG2025-09

用相似片段引导扩散模型,修复长段音乐缺失更自然。

Similarity-Guided Diffusion for Long-Gap Music Inpainting

  • 结合相似性检索与扩散模型,生成更连贯的音乐补全
  • 2秒缺失段修复时,主观评分优于无引导扩散和纯检索方法
  • 适合需要长序列音频修复的音乐生成与修复场景

音乐补全旨在重建受损录音中的缺失片段。尽管基于扩散的生成模型在中等长度缺失上表现良好,但在多秒级长间隙修复时往往难以保持音乐合理性。本文提出相似性引导扩散后验采样(SimDPS),一种融合扩散推理与相似性搜索的混合方法:先从语料库中检索上下文相似的候选片段,再将其融入修改后的似然函数,引导扩散过程生成上下文一致的重建结果。在钢琴音乐2秒缺失段的主观评估中,SimDPS相比无引导扩散显著提升感知合理性,且在候选片段中度相似时,性能常优于单独使用相似性搜索。结果表明,该混合相似性方法在长间隙扩散音频增强中具有潜力。

原文摘要 · Abstract (English)

Music inpainting aims to reconstruct missing segments of a corrupted recording. While diffusion-based generative models improve reconstruction for medium-length gaps, they often struggle to preserve musical plausibility over multi-second gaps. We introduce Similarity-Guided Diffusion Posterior Sampling (SimDPS), a hybrid method that combines diffusion-based inference with similarity search. Candidate segments are first retrieved from a corpus based on contextual similarity, then incorporated into a modified likelihood that guides the diffusion process toward contextually consistent reconstructions. Subjective evaluation on piano music inpainting with 2-s gaps shows that the proposed SimDPS method enhances perceptual plausibility compared to unguided diffusion and frequently outperforms similarity search alone when moderately similar candidates are available. These results demonstrate the potential of a hybrid similarity approach for diffusion-based audio enhancement with long gaps.

音乐生成扩散模型音频修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。