arXiv:2505.13032cs.SDcs.CL2025-05NeurIPS被引 133

构建跨音频模态的深度推理评测基准,推动听觉语言模型发展

MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

  • 设计四层推理框架(信号、感知、语义、文化),覆盖多模态音频场景
  • 包含1000个高质量音视频问答对,需多步深度推理,部分题需研究生级知识
  • 提供思维链标注,助力模型可解释性研究,适合音频与多模态方向学者

我们提出MMAR,一个用于评估音频-语言模型(ALMs)深度推理能力的新基准,涵盖大规模跨学科任务。该基准包含1,000个精心筛选的音频-问题-答案三元组,源自真实网络视频,并通过迭代纠错与质量检查确保数据质量。不同于仅限于特定声音、音乐或语音领域的现有基准,MMAR扩展至包含声音、音乐和语音混合的广泛现实音频场景。每个问题按四层推理层级分类:信号、感知、语义和文化,每层含子类别以体现任务多样性与复杂性。为促进研究,我们为每个问题标注了思维链(CoT)推理过程,推动音频推理进展。所有题目均需超越表层理解的多步深度推理,部分问题要求研究生级别感知与领域知识,显著提升难度。我们在多种模型上测试该基准,包括大音频-语言模型(LALMs)、大音频推理模型(LARMs)、全语言模型(OLMs)、大语言模型(LLMs)和大推理模型(LRMs),输入为音频描述。模型在MMAR上的表现揭示了当前模型在理解与推理能力上的关键局限,凸显该基准的挑战性。我们希望MMAR能成为推动这一重要但少被探索领域发展的催化剂。

原文摘要 · Abstract (English)

We introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,000 meticulously curated audio-question-answer triplets, collected from real-world internet videos and refined through iterative error corrections and quality checks to ensure high quality. Unlike existing benchmarks that are limited to specific domains of sound, music, or speech, MMAR extends them to a broad spectrum of real-world audio scenarios, including mixed-modality combinations of sound, music, and speech. Each question in MMAR is hierarchically categorized across four reasoning layers: Signal, Perception, Semantic, and Cultural, with additional sub-categories within each layer to reflect task diversity and complexity. To further foster research in this area, we annotate every question with a Chain-of-Thought (CoT) rationale to promote future advancements in audio reasoning. Each item in the benchmark demands multi-step deep reasoning beyond surface-level understanding. Moreover, a part of the questions requires graduate-level perceptual and domain-specific knowledge, elevating the benchmark's difficulty and depth. We evaluate MMAR using a broad set of models, including Large Audio-Language Models (LALMs), Large Audio Reasoning Models (LARMs), Omni Language Models (OLMs), Large Language Models (LLMs), and Large Reasoning Models (LRMs), with audio caption inputs. The performance of these models on MMAR highlights the benchmark's challenging nature, and our analysis further reveals critical limitations of understanding and reasoning capabilities among current models. We hope MMAR will serve as a catalyst for future advances in this important but little-explored area.

音频推理多模态评测基准思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。