arXiv:2507.04864cs.SDcs.LG2025-07中稿 · SMC 2025被引 1

用扩散模型复用音频样本,提升小数据训练效果并实现乐器替换。

Music Boomerang: Reusing Diffusion Models for Data Augmentation and Audio Manipulation

  • 基于预训练扩散模型,通过回旋采样复用现有音频。
  • 在小数据下提升节拍检测器性能,保留原节奏结构。
  • 可实现文本控制的单音轨乐器替换,适合音频增强任务。

生成式音乐音频模型通常仅依赖文本提示或旋律生成输出。最近提出的图像领域回旋采样(Boomerang sampling)可让任意预训练扩散模型生成接近已有样本的输出。本文探索其在音频领域的应用,作为数据增强或内容操控工具。具体在Stable Audio Open上实现回旋采样,用于增强一个前沿节拍追踪器的训练数据,并尝试替换录音中的乐器。结果表明,原有节奏结构基本保持不变,节拍追踪器性能有所提升,但仅在训练数据有限时有效;且可在单音轨输入下完成基于文本的乐器替换。我们开源了实现代码,鼓励在其他任务中探索数据增强及更多应用场景。

原文摘要 · Abstract (English)

Generative models of music audio are typically used to generate output based solely on a text prompt or melody. Boomerang sampling, recently proposed for the image domain, allows generating output close to an existing example, using any pretrained diffusion model. In this work, we explore its application in the audio domain as a tool for data augmentation or content manipulation. Specifically, implementing Boomerang sampling for Stable Audio Open, we augment training data for a state-of-the-art beat tracker, and attempt to replace musical instruments in recordings. Our results show that the rhythmic structure of existing examples is mostly preserved, that it improves performance of the beat tracker, but only in scenarios of limited training data, and that it can accomplish text-based instrument replacement on monophonic inputs. We publish our implementation to invite experiments on data augmentation in other tasks and explore further applications.

音频生成数据增强扩散模型乐器替换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。