arXiv:2606.22708cs.SDcs.AI2026-06被引 1

让大模型生成的乐谱具备可编辑的结构化特征。

Libretto: Giving LLM Agents a Sense of Musical Structure

论文配图:Libretto: Giving LLM Agents a Sense of Musical Structure
图 1 · 摘自论文原文
  • 用显式节拍、声部和小节组织的语法框架,让音乐可被语言模型理解。
  • 通过统计空间评估节奏、和声等六大维度,实现乐谱质量量化。
  • 适合音乐生成、教育创作与自动化修订场景,提升可控性。

生成式音乐系统如今能根据文本提示生成高质量音频,但输出难以检查、修改与诊断其音乐结构。我们提出 Libretto,一个面向语言模型代理的符号化音乐生成与修订框架。Libretto 采用以大模型为核心的语法结构,包含显式的起始时刻(onset)槽位、声部划分与小节级组织,并在基于语料校准的统计空间中评估每首作品在节奏、和声、旋律、织体、曲式与变奏六个维度的表现。相同的结构轴支持检索、诊断、抄袭风险控制与迭代自修订。在补全空白、参考引导的完整作曲、渐进变形及教育音乐生成任务中,Libretto 将符号化音乐从原始标记序列转变为可度量、可编辑的对象,使语言模型代理能高效操控音乐结构。

原文摘要 · Abstract (English)

Generative music systems can now produce impressive audio from text prompts, but audio outputs are difficult to inspect, edit, and diagnose as musical structure. We introduce Libretto, an agent-facing framework for symbolic music generation and revision. Libretto uses an LLM-native grammar with explicit onset slots, voices, and bar-level organization, then evaluates each piece in a corpus-calibrated statistical space over rhythm, harmony, melody, texture, form, and variation. The same structural axes support retrieval, diagnosis, copy-risk control, and iterative self-revision. Across gap filling, reference-guided full-piece generation, gradual morphing, and educational music generation, Libretto turns symbolic music from a raw token sequence into a measurable and editable object for language-model agents.

音乐生成大模型结构化符号化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。