arXiv:2606.31338cs.SDcs.AI2026-06

提出多维度诊断基准,检验音乐模型是否真懂乐器声音。

Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models

论文配图:Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models
图 1 · 摘自论文原文
  • 构建包含多种挑战的诊断测试集,覆盖不同音乐风格和时序场景。
  • 发现高准确率模型仍存在位置偏好与时间定位偏差问题。
  • 适合研究音频-语言对齐、音乐理解的学者与开发者参考。

近期音乐音频-语言模型在乐器问答基准上表现优异,但其性能是否反映真实的音频语义理解仍不明确。本文提出基于OpenMIC的诊断性评估序列,将二元乐器存在性问答扩展至减少流派先验的样本、易混淆乐器区分、更长音频上下文及时间定位任务。在这些设定下,高二元问答准确率常无法预测模型行为:模型表现出选项位置偏差、易混淆乐器误判以及时间响应偏差。结果表明,乐器理解应采用多维度诊断基准评估,而非单一准确率指标。

原文摘要 · Abstract (English)

Recent music audio-language models achieve high accuracy on instrument question-answering benchmarks, but it remains unclear whether this reflects robust audio grounding or benchmark-specific shortcuts. In this paper, we introduce an OpenMIC-derived diagnostic benchmark sequence for instrument grounding in music audio-language models, extending binary instrument-presence QA to genre-prior-reduced examples, confusable instrument discrimination, longer audio context, and temporal localization. Across these settings, high binary QA accuracy often fails to predict model behavior: models can exhibit option-position bias, confusable-instrument errors, and temporal response bias. These results suggest that instrument grounding should be evaluated with multi-axis diagnostic benchmarks rather than a single aggregate accuracy.

音乐理解音频语言模型诊断评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。