arXiv:2608.03999cs.SDcs.CL2026-08

音乐标记方式比模型大小更影响生成质量,性能级标记显著提升音符生成效果。

Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation

论文配图:Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation
图 1 · 摘自论文原文
  • 采用性能级音乐标记(PMT),精确到10毫秒时间与力度,实现高保真音符表达。
  • 0.8B模型用PMT时FMD达159,远低于27B模型在节拍网格下的286,差距最大达2.8倍。
  • 效果稳定可复现,且不依赖特定模型或数据集,适合关注音乐生成质量的研究者。

文本转音乐的语言模型通常默认使用某种音乐标记方式,但其独立影响从未被测量。我们固定Qwen3.5(0.8B-27B)模型、数据、预算与解码策略,仅替换七种不同标记方式,以各表示的无模型上限为基准。结果清晰且出人意料:表示方式而非模型规模是分布保真度的关键变量。模型扩大34倍,弗雷切特音乐距离(FMD)几乎不变;而更换表示方式,FMD减半。我们提出的性能分辨率标记(PMT,10ms精度,每音符力度,多轨纹理;609个符号)在0.8B模型上达到FMD 159,显著优于节拍网格(FMD 272–286,降低1.7–1.8倍,局部高达2.8倍;置信区间不重叠)。0.8B PMT模型甚至超越27B节拍网格模型。该优势在2600万条从零训练的骨干网络和另一款性能标记中重现,表明属于类别特性。即便将PMT对齐至节拍网格分辨率,仍领先67–129 FMD。该效应为分布层面的,是否可听仍待人类评估验证(已预注册研究)。轻量级解码约束可使乐器识别F1从0.28升至0.60,调性正确率从0.16升至0.35,且无分布代价。我们发布评估框架、25+检查点、两个语料库(共86.6k对齐样本,含歌词/音频/ABC/MIDI;625万标注文本,为最大音乐语料),以及印刻诊断工具:现有文本转MIDI系统在不同输入下仍保持近似训练分布(72%对比71%和弦-时间匹配)。今后的表示主张可被实证衡量,而非仅宣称。

原文摘要 · Abstract (English)

Text-to-music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe, its effect has never been measured in isolation. We fix pretrained Qwen3.5 (0.8B-27B), data, budget, and decoding, and swap only the representation across seven tokenizations, anchoring texture metrics to each representation's model-free ceiling. The ordering is clean and surprising: representation, not model size, is the binding variable for distributional fidelity. Scaling the backbone 34x barely moves Frechet Music Distance (FMD), whereas switching representation halves it. PMT, a performance-resolution stream we release (10 ms timing, per-note velocity, multi-track texture; 609 symbols), reaches FMD 159 at 0.8B against 272-286 for beat grids (1.7-1.8x lower, up to 2.8x elsewhere; non-overlapping bootstrap CIs), so a 0.8B performance-resolution model beats a 27B beat grid. It reappears on a 26M from-scratch backbone and a second performance-resolution tokenizer: a property of the class, not one lucky vocabulary. Nor is it a finer-lattice artifact: snapping PMT's onsets to the beat grids' resolution still leaves it 67-129 FMD ahead of both (n=500). The effect is distributional; whether it is audible is a separate question, left open by our probe, with a human study pre-registered. Native caption adherence is weak but separable: a lightweight decode-time constraint doubles instrument-F1 (.28 to .60) and Correct-Key (.16 to .35) at no distributional cost. We release the harness, 25+ checkpoints, two corpora (86.6k aligned across caption/MIDI/ABC/audio; 6.25M captioned, the largest for music), and an imprinting diagnostic: published text-to-MIDI systems reproduce their training distribution near-invariant to the caption (72% vs. 71% chord-time on disjoint domains). The field's next representation claim can now be measured, not asserted.

音乐生成标记方法文本转音乐扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。