arXiv:2509.21714cs.SDcs.MM2025-09

让音乐生成可编程编辑,改一段不重跑整首歌

MusicWeaver: Programmable Long-Form Music Generation with Provably Local Editing

  • 用程序层控制音乐结构,实现意图与音频的精准对接
  • 支持局部修改且不破坏整体一致性,连续编辑质量稳定
  • 能从录音逆向生成可编辑程序,适合创作者和交互设计

音乐生成系统虽日益逼真,但创作者难以直接控制结构,修改需重生成整首曲。本文提出 MusicWeaver,将音乐创作重构为程序引导生成:规划阶段生成多层级歌曲程序,包含乐段、动机重复与小节属性;渲染阶段据此合成音频。定义了一套类型化计划操作(替换、插入、删除、重标记、属性修改),并证明其保持计划有效性。提出投影扩散修复方法,确保修改区域外音频完全不变。设计全局-局部扩散变换器与动机记忆检索机制,可执行分钟级计划并生成一致又多样的段落返回。系统还支持从录音中恢复可编辑程序,以及自然语言指令编译。引入计划忠实度与编辑保真度指标,经人工评估验证。在文本、视频及二者结合条件下的生成实验中,结构连贯性与可编辑性达当前最优,且连续编辑质量不下降。

原文摘要 · Abstract (English)

Music generation systems produce increasingly realistic audio, yet they expose no interface between a creator's structural intent and the rendered sound. Form can only be steered through prompts, while local revisions require regenerating entire pieces. We recast music creation as program-guided generation, using an explicit, human-interpretable program layer between intent and audio. We present MusicWeaver, which decomposes generation into planning and rendering. The planning stage predicts a structured plan, a multi-level song program encoding musical form, motif recurrence, and bar-level attributes. The rendering stage synthesizes audio conditioned on this plan. We formalize editing as an algebra of typed plan operations, including section replacement, insertion, deletion, motif retagging, and attribute changes, and prove that they preserve plan validity. To realize edits, we propose Projected Diffusion Inpainting, which guarantees that audio outside the edited span is preserved exactly across repeated revisions. For rendering, we design a Global-Local Diffusion Transformer with Motif Memory Retrieval that executes minute-scale plans and produces section returns that are consistent yet varied. Our framework also supports plan induction, recovering an editable program from an existing recording, and a natural-language editor that compiles free-form instructions into validated operation sequences. We introduce plan-faithfulness and edit-fidelity measures and validate them against human judgments. Experiments on generation conditioned on text, video, and their combination show state-of-the-art structural coherence and editability, while edit quality remains stable over sequential revisions where prior methods degrade.

音乐生成可编辑性扩散模型程序生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。