通过优化生成时间步,实现音乐中乐器的精准编辑。
Diff-TONE: Timestep Optimization for iNstrument Editing in Text-to-Music Diffusion Models
- 利用分类器选择合适中间时间步,控制乐器替换。
- 编辑后保留原曲内容,同时准确改变音色特征。
- 无需训练模型,保持生成速度,适合音乐创作应用。
文本到音乐生成模型的突破正在重塑创作生态,为作曲家提供前所未有的创作与实验工具。然而,精确控制生成过程以达到特定目标仍具挑战性:即使提示词微调且随机种子相同,生成结果也可能大幅变化。本文探索现有文本到音乐扩散模型在乐器编辑中的应用,旨在对已有音频轨道进行乐器替换,同时保留原始内容。基于模型先关注整体结构、再添加乐器信息、最后优化质量的生成规律,我们发现通过乐器分类器识别的合适中间时间步,可平衡内容保真与音色变换效果。该方法无需额外训练模型,也不影响生成速度。
原文摘要 · Abstract (English)
Breakthroughs in text-to-music generation models are transforming the creative landscape, equipping musicians with innovative tools for composition and experimentation like never before. However, controlling the generation process to achieve a specific desired outcome remains a significant challenge. Even a minor change in the text prompt, combined with the same random seed, can drastically alter the generated piece. In this paper, we explore the application of existing text-to-music diffusion models for instrument editing. Specifically, for an existing audio track, we aim to leverage a pretrained text-to-music diffusion model to edit the instrument while preserving the underlying content. Based on the insight that the model first focuses on the overall structure or content of the audio, then adds instrument information, and finally refines the quality, we show that selecting a well-chosen intermediate timestep, identified through an instrument classifier, yields a balance between preserving the original piece's content and achieving the desired timbre. Our method does not require additional training of the text-to-music diffusion model, nor does it compromise the generation process's speed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。