音乐标记方式比模型大小更影响生成质量,性能级标记显著提升音符生成效果。
Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation

- 采用性能级音乐标记(PMT),精确到10毫秒时间与力度,实现高保真音符表达。
- 0.8B模型用PMT时FMD达159,远低于27B模型在节拍网格下的286,差距最大达2.8倍。
- 效果稳定可复现,且不依赖特定模型或数据集,适合关注音乐生成质量的研究者。
文本转音乐的语言模型通常默认使用某种音乐标记方式,但其独立影响从未被测量。我们固定Qwen3.5(0.8B-27B)模型、数据、预算与解码策略,仅替换七种不同标记方式,以各表示的无模型上限为基准。结果清晰且出人意料:表示方式而非模型规模是分布保真度的关键变量。模型扩大34倍,弗雷切特音乐距离(FMD)几乎不变;而更换表示方式,FMD减半。我们提出的性能分辨率标记(PMT,10ms精度,每音符力度,多轨纹理;609个符号)在0.8B模型上达到FMD 159,显著优于节拍网格(FMD 272–286,降低1.7–1.8倍,局部高达2.8倍;置信区间不重叠)。0.8B PMT模型甚至超越27B节拍网格模型。该优势在2600万条从零训练的骨干网络和另一款性能标记中重现,表明属于类别特性。即便将PMT对齐至节拍网格分辨率,仍领先67–129 FMD。该效应为分布层面的,是否可听仍待人类评估验证(已预注册研究)。轻量级解码约束可使乐器识别F1从0.28升至0.60,调性正确率从0.16升至0.35,且无分布代价。我们发布评估框架、25+检查点、两个语料库(共86.6k对齐样本,含歌词/音频/ABC/MIDI;625万标注文本,为最大音乐语料),以及印刻诊断工具:现有文本转MIDI系统在不同输入下仍保持近似训练分布(72%对比71%和弦-时间匹配)。今后的表示主张可被实证衡量,而非仅宣称。
原文摘要 · Abstract (English)
Text-to-music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe, its effect has never been measured in isolation. We fix pretrained Qwen3.5 (0.8B-27B), data, budget, and decoding, and swap only the representation across seven tokenizations, anchoring texture metrics to each representation's model-free ceiling. The ordering is clean and surprising: representation, not model size, is the binding variable for distributional fidelity. Scaling the backbone 34x barely moves Frechet Music Distance (FMD), whereas switching representation halves it. PMT, a performance-resolution stream we release (10 ms timing, per-note velocity, multi-track texture; 609 symbols), reaches FMD 159 at 0.8B against 272-286 for beat grids (1.7-1.8x lower, up to 2.8x elsewhere; non-overlapping bootstrap CIs), so a 0.8B performance-resolution model beats a 27B beat grid. It reappears on a 26M from-scratch backbone and a second performance-resolution tokenizer: a property of the class, not one lucky vocabulary. Nor is it a finer-lattice artifact: snapping PMT's onsets to the beat grids' resolution still leaves it 67-129 FMD ahead of both (n=500). The effect is distributional; whether it is audible is a separate question, left open by our probe, with a human study pre-registered. Native caption adherence is weak but separable: a lightweight decode-time constraint doubles instrument-F1 (.28 to .60) and Correct-Key (.16 to .35) at no distributional cost. We release the harness, 25+ checkpoints, two corpora (86.6k aligned across caption/MIDI/ABC/audio; 6.25M captioned, the largest for music), and an imprinting diagnostic: published text-to-MIDI systems reproduce their training distribution near-invariant to the caption (72% vs. 71% chord-time on disjoint domains). The field's next representation claim can now be measured, not asserted.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。