构建音乐后期制作的智能代理框架,让模型能像真人一样精细修改音乐。
RIME: Enabling Large-Scale Agentic Music Post-Production

- 基于真实混音流程生成带指令的音乐修改数据对
- 创建3000组指令与真实音频配对数据,验证模型能力短板
- 适合想用AI辅助音乐制作的创作者和研究者
几乎每首你听过的录音音乐都经过后期修改;商业发行极少是音乐人直接输出的成品。尽管音乐生成模型可实现一次性输出,但精细的迭代优化流程仍是其难以触及的领域。同时,音乐人虽能表达想要的效果,却未必掌握复杂制作工具。本文将此问题形式化为「智能后期制作」任务,即针对歌曲各部分进行定向优化并整合成最终作品。我们指出瓶颈在于数据:现有语料库无法反映真实后期流程中音乐人与工程师实际使用的术语体系。我们提出规则驱动的音乐编辑指令(RIME)框架,从任意基础音乐数据集出发,基于标准制作方法、设计模式和约束生成真实可信的指令-音频配对数据。RIME结合新工具POEMS,实现分轨分离、混音与常用效果处理,供多模态智能体使用。我们利用RIME和POEMS生成3000对编辑指令与真实音频,并以此评估现有多模态大模型在该任务上的表现,揭示当前模型在后期制作能力上的持续不足。此外,我们证明了通过监督微调可显著提升后制智能体性能。RIME被视为迈向迭代式音乐智能体的关键一步,未来有望像交互式编程代理重塑软件工程那样,改变音乐生产方式。
原文摘要 · Abstract (English)
Almost every piece of recorded music you have ever heard was modified before it reached you; commercial releases rarely spring fully-formed from the mind of a musician. Despite the promise of music generation models for one-shot output, such fine-grained iterative refinement workflows are a complementary problem, and largely out of their reach. There is also a gap for musicians: while they can express what they want to hear, not all have the facility with studio production tools to implement the complex set of actions needed to realize these intuitions. We formalize this task as agentic post-production, wherein individual aspects of a song are targeted, refined, and combined into a final track. We argue the bottleneck is data: existing corpora do not reflect how realistic post-production chains map onto the vocabulary musicians and engineers actually use. We argue there is a language for modifying recorded music that is dense, consistent, and learnable. We introduce the Rule-based Instructions for Music Editing (RIME) framework, which generates realistic paired edit-instruction data from any baseline music dataset grounded in canonical methods, design patterns, and constraints derived from real production workflows. RIME leverages POEMS, a new toolkit that combines stem separation, mixing, and common studio effects for use by multimodal agents. We use RIME and POEMS to generate 3,000 pairs of edit instructions and ground truth audio, and use this data to evaluate existing multimodal LLMs as agents on this task, showing persistent challenges in current models' post-production capabilities. We also demonstrate RIME's ability to improve post-production agent performance via supervised fine-tuning. We see RIME as an early step toward iterative musical agents, collaborative systems that could transform music production much as interactive coding agents have reshaped software engineering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。