MusRec实现零样本文本音乐编辑,无需重训即可高效修改真实音乐。
MusRec: Zero-Shot Text-to-Music Editing via Rectified Flow and Diffusion Transformers
- 结合修正流与扩散Transformer,直接编辑真实音乐
- 在内容保留和结构一致性上优于现有方法
- 适合游戏影视配乐与个性化音乐定制场景
音乐编辑已成为人工智能的重要实用领域,应用于游戏、电影配乐及根据用户偏好个性化现有曲目。然而,现有模型存在显著局限:仅能编辑自身生成的合成音乐、需高精度提示或任务特定重训练,缺乏真正的零样本能力。本文利用修正流与扩散Transformer的最新进展,提出MusRec——一种可在真实世界音乐上高效执行多样编辑任务的零样本文本到音乐编辑模型。实验表明,该方法在保持音乐内容、结构一致性和编辑保真度方面均优于现有方法,为真实场景下的可控音乐编辑奠定了坚实基础。
原文摘要 · Abstract (English)
Music editing has emerged as an important and practical area of artificial intelligence, with applications ranging from video game and film music production to personalizing existing tracks according to user preferences. However, existing models face significant limitations, such as being restricted to editing synthesized music generated by their own models, requiring highly precise prompts, or necessitating task-specific retraining, thus lacking true zero-shot capability. leveraging recent advances in rectified flow and diffusion transformers, we introduce MusRec, a zero-shot text-to-music editing model capable of performing diverse editing tasks on real-world music efficiently and effectively. Experimental results demonstrate that our approach outperforms existing methods in preserving musical content, structural consistency, and editing fidelity, establishing a strong foundation for controllable music editing in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。