自动构建音乐感知测评框架,支持多模态灵活评估大模型音乐理解能力。
Music I Care About: Automated Multimodal Benchmarking of LLM Music Perception Skills on (Almost) Any Music

- 基于用户数据自动生成多模态测评题,融合乐谱、音频等格式。
- 在ChoraleBricks数据集上验证,小规模基准即可实现可靠对比结果。
- 适用于研究者快速评估模型音乐感知能力,尤其适合跨模态分析场景。
音乐是人类文化的核心,以音频、符号表示(如MIDI、MusicXML)和乐谱图像等多种模态存在。尽管多模态大语言模型(MLLMs)快速发展,现有音乐测评基准仍存在三大局限:一是静态大规模基准计算成本高,结果泛化性未知;二是所谓“音乐理解”测评常未真正考察音乐感知;三是无法系统比较不同模态的表现。为此,我们提出音乐感知元基准框架MusICA-MetaBench,可基于用户提供的数据自动构建按需测评集。通过结构化符号表示(如MusicXML)与预设问题模板,生成涵盖音频、乐谱图像与符号文件的多项选择题,探测音乐感知能力,符合音乐教学逻辑。我们在ChoraleBricks数据集上验证框架有效性,并实验确定了确保统计可靠性所需的基准规模。通过与文本仅和白噪声基线对比,证明所提问题确能衡量音乐感知。MusICA-MetaBench为MLLM音乐感知的跨模态评估提供了新范式,实现高效、定制化的评估能力。
原文摘要 · Abstract (English)
Music represents a cornerstone of human culture, existing digitally across diverse modalities, including audio, symbolic encodings (e.g., MIDI, MusicXML), and sheet music. Despite the advancement of Multimodal Large Language Models (MLLMs), current music benchmarks face three major limitations. First, large static benchmarks are resource-intensive to evaluate, and it remains unclear how their results transfer to diverse kinds of music beyond those included in the benchmark. Second, benchmarks claiming to measure "music understanding" often fail to require music perception. Third, they do not support systematic performance comparisons across musical modalities. To overcome these issues, we introduce the Music I Care About Meta-Benchmark (MusICA-MetaBench), a framework that automatically derives on-demand benchmarks directly from user-provided data. By leveraging structured symbolic representations (e.g., MusicXML) and our pre-defined question templates, we build multiple-choice question-answer pairs that probe music perception competencies, aligned with music pedagogy, across audio, music notation images, and symbolic files. We demonstrate our framework with the ChoraleBricks dataset, and experimentally determine benchmark sizes that ensure statistically reliable model comparisons for this setup. By comparing against text-only and white-noise baselines, we show our questions do measure music perception. Ultimately, MusICA-MetaBench represents a significant advancement in the cross-modal assessment of music perception for MLLMs. By proposing a dataset-specific benchmarking paradigm, it enables efficient on-demand evaluation of music perception capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。