首个真实音乐场景下的多模态音乐推理评测基准
WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning
- 基于真实乐谱与用户问题构建评测数据集
- 揭示主流多模态大模型在音乐推理中的强弱项
- 适合音乐人工智能与跨模态研究者参考
近年来,多模态大语言模型(MLLMs)在视觉-语言任务中展现出卓越能力,但在多模态符号音乐领域的推理能力仍鲜被探索。本文提出WildScore,首个面向真实世界音乐评分的多模态符号音乐推理与分析基准,用于评估MLLMs理解实际乐谱并回答复杂音乐学问题的能力。每个数据实例均来自真实音乐作品,附带真实的用户生成问题与讨论,捕捉实际音乐分析的复杂性。为实现系统化评估,我们构建了包含高层与细粒度音乐学本体的分类体系,并将复杂音乐推理建模为多项选择题问答任务,支持可控且可扩展的评估。对前沿MLLMs在WildScore上的实证测试揭示了其在视觉-符号推理中的有趣模式,指明了符号音乐推理中的潜力方向与持续挑战。数据集与代码已开源。
原文摘要 · Abstract (English)
Recent advances in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities across various vision-language tasks. However, their reasoning abilities in the multimodal symbolic music domain remain largely unexplored. We introduce WildScore, the first in-the-wild multimodal symbolic music reasoning and analysis benchmark, designed to evaluate MLLMs' capacity to interpret real-world music scores and answer complex musicological queries. Each instance in WildScore is sourced from genuine musical compositions and accompanied by authentic user-generated questions and discussions, capturing the intricacies of practical music analysis. To facilitate systematic evaluation, we propose a systematic taxonomy, comprising both high-level and fine-grained musicological ontologies. Furthermore, we frame complex music reasoning as multiple-choice question answering, enabling controlled and scalable assessment of MLLMs' symbolic music understanding. Empirical benchmarking of state-of-the-art MLLMs on WildScore reveals intriguing patterns in their visual-symbolic reasoning, uncovering both promising directions and persistent challenges for MLLMs in symbolic music reasoning and analysis. We release the dataset and code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。