构建音乐理解基准,揭示大模型在听觉关系推理上的显著短板。
The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS
- 设计10项任务评估音乐感知与关系推理能力
- 4个主流模型表现参差,部分接近随机水平
- 提示词链式推理反而降低性能,适合研究音频模型的开发者
多模态大语言模型在音频理解方面已展现能力,但现有评估可能掩盖其在关系推理方面的根本缺陷。我们提出Music Understanding and Structural Evaluation(MUSE)基准,一个开源资源,包含10项任务,用于探测基础音乐感知技能。我们对四个顶级模型(Gemini Pro和Flash、Qwen2.5-Omni、Audio-Flamingo 3)进行了评估,并与大规模人类基线(N=200)对比。结果揭示了当前SOTA模型能力差异显著,且与人类专家存在持续差距。尽管Gemini Pro在基础感知任务上表现良好,但Qwen和Audio Flamingo 3的表现接近随机水平,暴露出严重的感知缺陷。此外,我们发现链式思维(CoT)提示效果不稳定,常产生负面结果。本工作为评估音乐表征的不变性提供了关键工具,推动更鲁棒的AI系统发展。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have demonstrated capabilities in audio understanding, but current evaluations may obscure fundamental weaknesses in relational reasoning. We introduce the Music Understanding and Structural Evaluation (MUSE) Benchmark, an open-source resource with 10 tasks designed to probe fundamental music perception skills. We evaluate four SOTA models (Gemini Pro and Flash, Qwen2.5-Omni, and Audio-Flamingo 3) against a large human baseline (N=200). Our results reveal a wide variance in SOTA capabilities and a persistent gap with human experts. While Gemini Pro succeeds on basic perception, Qwen and Audio Flamingo 3 perform at or near chance, exposing severe perceptual deficits. Furthermore, we find Chain-of-Thought (CoT) prompting provides inconsistent, often detrimental results. Our work provides a critical tool for evaluating invariant musical representations and driving development of more robust AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。