arXiv:2603.27877cs.CLcs.SD2026-03被引 2

构建首个人工撰写音乐理解问答数据集,评测大模型真实听懂音乐的能力。

HumMusQA: A Human-written Music Understanding QA Benchmark Dataset

  • 320道专家手写问题,覆盖音乐感知与理解深层维度
  • 6个顶尖音频语言模型在新数据集上表现参差,最高准确率仅61.8%
  • 揭示模型依赖单一模态捷径的脆弱性,适合评估多模态理解能力

大型音频-语言模型(LALMs)在音乐理解方面的评估需要一个严谨定义的基准,以真正检验模型是否能感知和解释音乐——而当前的数据方法常无法满足这一标准。本文提出一种精心设计的音乐评估框架,引入一个新的基准数据集HumMusQA,包含320道由具音乐训练背景的专家手工编写并验证的问题,论证了这种专注的、人工标注方式在探测复杂音频理解方面的优势。为展示该数据集的应用,我们对六种最先进的LALMs进行了基准测试,并进一步检验其对单模态捷径的鲁棒性。

原文摘要 · Abstract (English)

The evaluation of music understanding in Large Audio-Language Models (LALMs) requires a rigorously defined benchmark that truly tests whether models can perceive and interpret music, a standard that current data methodologies frequently fail to meet. This paper introduces a meticulously structured approach to music evaluation, proposing a new dataset of 320 hand-written questions curated and validated by experts with musical training, arguing that such focused, manual curation is superior for probing complex audio comprehension. To demonstrate the use of the dataset, we benchmark six state-of-the-art LALMs and additionally test their robustness to uni-modal shortcuts.

音乐理解多模态评测基准LALM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。