测试大模型对音乐符号语言的理解与生成能力
Can LLMs understand LilyPond? A benchmark for symbolic music generation and understanding
- 基于LilyPond构建统一评估基准,覆盖生成与理解任务
- 零样本下可生成可执行乐谱,但结构理解仍困难
- 建议多指标联合评估,避免单一分数误导
大型语言模型在符号化音乐生成与理解方面的评估仍分散于不同表示、数据集和度量方式。我们提出LilyBench,一个基于LilyPond的基准,对同一组开源大模型进行生成与理解的联合评估。该基准包含200个提示的生成任务和10项从ABC-Eval改编的理解任务,涵盖语法、元数据预测、结构序列和音乐识别。生成质量通过编译率、MusPy描述符分布(使用Jensen-Shannon相似性)以及基于LilyBERT的Fréchet音乐距离(FMD)评估。在四个开源模型上的实验表明,零样本条件下可实现可执行的LilyPond生成,但结构理解任务仍具挑战性,尽管在作曲家和流派识别上表现良好。实验还揭示了描述符与嵌入式度量之间的系统性分歧,提示符号音乐评估应采用多指标三角验证而非单一分数排名。我们已公开基准、提示库与评估代码,支持未来研究。
原文摘要 · Abstract (English)
Symbolic music evaluation for large language models remains fragmented across representations, datasets, and metrics. We introduce LilyBench, a LilyPond-based benchmark that jointly evaluates symbolic music generation and music understanding on the same family of open-weight LLMs. The benchmark includes a 200-prompt generation suite and ten understanding tasks adapted from ABC-Eval, covering syntax, metadata prediction, structural sequencing, and music recognition. Generation quality is evaluated using compile rate, MusPy descriptor distributions via Jensen-Shannon similarity, and LilyBERT-based Fréchet Music Distance (FMD). Experiments on four open-weight models show that executable LilyPond generation is achievable in zero-shot settings, while structural understanding tasks remain challenging despite strong performance on composer and genre recognition. Our experiments also reveal systematic disagreements between descriptor-based and embedding-based metrics, suggesting that symbolic music evaluation benefits from metric triangulation rather than single-score ranking. We release the benchmark, prompt bank, and evaluation code to support future research in symbolic music generation and understanding at https://github.com/CSCPadova/lilybench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。