评测模型理解乐谱与演奏的跨模态能力,发现大模型表现仍差。
MuSP-Bench: Advanced Multimodal Benchmarking of Music Understanding across Score and Performance
- 构建490道人工设计题目,覆盖乐谱、演奏、诠释与长时推理
- 模型在乐谱理解上表现不佳,演奏音频推理更弱
- 适合音乐信息检索与跨模态理解研究者参考
音乐家通常通过乐谱和演奏来传达音乐。乐谱表达音乐意图,演奏则以声音实现它。为探究模型能否有效理解这两种模态,我们提出MuSP-Bench,一个包含490道人工设计问题的基准测试,聚焦于乐谱与演奏之间的跨模态理解。该基准涵盖古典钢琴与管弦乐作品,覆盖基于乐谱、基于演奏、诠释性及长时程推理任务。我们在多种输入条件下评估前沿多模态大模型。结果表明,这些模型在理解乐谱方面已存在明显短板,而在分析演奏音频时面临更大挑战。该基准可访问:https://musp.vaclis.net/。
原文摘要 · Abstract (English)
Musicians commonly communicate music through scores and performances. Scores encode musical intent, while performances realize it in sound. To investigate whether models can meaningfully engage with both modalities, we introduce MuSP-Bench, a human-authored benchmark of 490 questions targeting understanding across Musical Scores and Performances. The benchmark distinguishes itself by spanning score-based, performance-based, interpretive, and long-horizon reasoning across classical piano and orchestral works. We evaluate frontier multimodal large language models under multiple input conditions. Our results show that these models struggle substantially to understand scores, while facing even greater challenges when reasoning about performance audio. The benchmark is available at https://musp.vaclis.net/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。