用分数蒸馏提升文本引导音乐编辑的一致性与精度
SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing
- 通过增量去噪分数实现零样本音乐编辑
- 保留原始音乐内容,编辑准确率显著提升
- 支持个性化风格编辑,适合音乐创作与影视配乐
音乐编辑在音乐制作中至关重要,广泛应用于游戏开发与影视制作。现有零样本文本引导编辑方法多依赖预训练扩散模型的前向-后向扩散过程,但常难以保持音乐内容一致性。此外,仅靠文本指令往往无法准确描述目标音乐。本文提出两种改进方法:SteerMusic采用粗粒度零样本编辑,利用增量去噪分数;SteerMusic+则通过操控用户定义的风格概念令牌,实现细粒度个性化音乐编辑,可生成仅靠文本无法表达的音乐风格。实验表明,所提方法在保持音乐内容一致性和编辑保真度方面优于现有方法。用户研究进一步验证了其卓越的编辑质量。
原文摘要 · Abstract (English)
Music editing is an important step in music production, which has broad applications, including game development and film production. Most existing zero-shot text-guided editing methods rely on pretrained diffusion models by involving forward-backward diffusion processes. However, these methods often struggle to preserve the musical content. Additionally, text instructions alone usually fail to accurately describe the desired music. In this paper, we propose two music editing methods that improve the consistency between the original and edited music by leveraging score distillation. The first method, SteerMusic, is a coarse-grained zero-shot editing approach using delta denoising score. The second method, SteerMusic+, enables fine-grained personalized music editing by manipulating a concept token that represents a user-defined musical style. SteerMusic+ allows for the editing of music into user-defined musical styles that cannot be achieved by the text instructions alone. Experimental results show that our methods outperform existing approaches in preserving both music content consistency and editing fidelity. User studies further validate that our methods achieve superior music editing quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。