arXiv:2511.20697cs.SDcs.AI2025-11ACL被引 6

评测大模型理解完整乐谱的能力,发现视觉与文本模态差距明显。

Musical Score Understanding Benchmark: Evaluating Large Language Models' Comprehension of Complete Musical Scores

  • 构建跨模态乐谱理解基准,涵盖巴赫到德彪西等作曲家作品。
  • 超过15个模型测试显示,微调后性能显著提升但多层级一致性仍难保证。
  • 适合研究多模态推理、音乐AI理解能力的学者使用。

理解完整的乐谱需要对音高、节奏、和声及宏观结构进行综合推理,但目前大型语言模型与视觉-语言模型在完整乐谱解读方面的能力尚未充分评估。本文提出音乐乐谱理解基准(MSU-Bench),一个由人工标注的跨模态乐谱理解基准,支持文本(ABC记谱法)与视觉(PDF)两种形式。该基准包含来自巴赫、贝多芬、肖邦、德彪西等作曲家作品的1,800组生成式问答对,按难度分为四个等级,从音符起始信息到织体与曲式分析逐步递进。对15种以上先进模型在零样本与微调设置下的评估表明,不同模态间存在显著差距,层级表现不稳定,维持多层级正确性仍具挑战。微调可显著提升各模态表现,同时保持通用知识,使MSU-Bench成为未来多模态推理研究的可靠基础。数据集与代码已开源:https://github.com/Congren-Dai/MSU-Bench。

原文摘要 · Abstract (English)

Understanding complete musical scores entails integrated reasoning over pitch, rhythm, harmony, and large-scale structure, yet the ability of Large Language Models and Vision--Language Models to interpret full musical notation remains insufficiently examined. We introduce Musical Score Understanding Benchmark (MSU-Bench), a human-curated benchmark for score-level musical understanding across textual (ABC notation) and visual (PDF) modalities. MSU-Bench contains 1,800 generative question-answer pairs from works by Bach, Beethoven, Chopin, Debussy, and others, organised into four levels of increasing difficulty, ranging from onset information to texture and form. Evaluations of more than fifteen state-of-the-art models, in both zero-shot and fine-tuned settings, reveal pronounced modality gaps, unstable level-wise performance, and challenges in maintaining multilevel correctness. Fine-tuning substantially improves results across modalities while preserving general knowledge, positioning MSU-Bench as a robust foundation for future research in multimodal reasoning. The benchmark and code are available at https://github.com/Congren-Dai/MSU-Bench.

乐谱理解多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。