用自然语言指令零样本编辑鼓点,让AI听懂音乐创作意图。
Not that Groove: Zero-Shot Symbolic Music Editing
- 设计文本化鼓点符号系统,让大模型理解音乐逻辑并推理编辑
- 在自动化测试中实现68%的编辑成功率,接近专业音乐人水平
- 适合音乐制作人、创作者,无需标注数据即可实现精准控制
尽管近年来人工智能音乐生成多聚焦于直接音频合成,但这类系统存在固有僵化,难以满足专业音乐制作人对精细可控创作的需求。符号化音乐(如MIDI)通过可编辑的音符级参数解决此问题,然而基于指令的符号音乐编辑因缺乏成对的指令-MIDI数据集而进展缓慢。本文提出将零样本符号音乐编辑形式化为结构化推理任务,引入一种基于文本的“drumroll”记谱法,将音乐机制转化为空间化、语法驱动的网格,使现成的大语言模型(LLMs)仅通过零样本提示即可逻辑推导并应用复杂鼓点修改。为严格评估该范式,我们构建了Not that Groove基准,包含数千个鼓点及其具体、描述性与风格化的自然语言指令。关键在于,为克服人工评价成本高且主观的问题,我们提出一种可扩展的、领域知情的自动化单元测试框架,通过符号验证编辑后的鼓点是否满足用户请求的核心约束。在八种先进LLMs上的实验表明,最优模型在自动化测试中达到68%的成功率。听觉测试进一步证实,该程序化测试与专业音乐人主观判断高度一致,建立了高效、可扩展、数据友好的可控AI音乐生产新基础。
原文摘要 · Abstract (English)
While recent advancements in AI music generation have predominantly focused on direct audio synthesis, these systems suffer from inherent rigidity, limiting their utility for professional music producers who require granular, highly malleable creative control. Symbolic music (e.g., MIDI) resolves this constraint by providing editable note-level parameters, yet the natural progression to instruction-driven symbolic music editing remains critically under-explored due to a severe scarcity of paired instruction-MIDI datasets. In this paper, we bypass this data bottleneck by formalizing zero-shot symbolic music editing as a structured reasoning task. We introduce a novel text-based "drumroll" notation that translates musical mechanics into a spatial, syntax-driven grid, empowering off-the-shelf Large Language Models (LLMs) to logically deduce and apply complex edits to drum grooves using only zero-shot prompting. To rigorously evaluate this paradigm, we propose Not that Groove, a comprehensive benchmark comprising thousands of drum grooves paired with specific, descriptive, and stylistic natural language instructions. Crucially, to overcome the prohibitive cost and subjectivity of human musical evaluation, we introduce a scalable, domain-informed automated unit-testing framework that symbolically verifies whether an edited groove satisfies the core constraints of the user's request. Our extensive experiments across eight state-of-the-art LLMs demonstrate the high efficacy of this approach, with the top-performing model achieving a 68% success rate on our automated unit tests. Furthermore, listening tests confirm that our programmatic unit tests align highly with the subjective judgments of professional musicians, establishing a robust, data-efficient, and scalable foundation for the future of controllable AI music production.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。