构建首个全面评估音频通用智能的基准,挑战当前AI在听觉理解上的极限。
MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
- 设计5305个真实场景音频问答对,覆盖49种听觉技能与多模态推理
- 顶尖模型如Gemini 2.5 Flash仅达59.2%准确率,多项任务接近随机水平
- 专为长时音频、空间推理等复杂能力设计,适合研究通用人工智能者
语音、非语音声音和音乐的理解是实现人类级智能的关键。因此,人工智能代理必须展现出全面的音频理解能力,才能被视为具备通用智能。然而,全面评估听觉智能仍具挑战。为此,我们提出MMAU-Pro,目前最全面且严格筛选的评估人工智能听觉智能的基准。MMAU-Pro包含5,305个实例,每个实例配有由人类专家生成的问答对,涵盖语音、声音、音乐及其组合。不同于现有基准,MMAU-Pro在49种独特技能及多个复杂维度上评估听觉智能,包括长时音频理解、空间音频推理、多音频理解等。所有问题均需经过多步推理,采用选择题与开放回答两种格式。重要的是,音频数据直接来自真实世界('from the wild'),而非已有数据集中的已知分布。我们评估了22个领先的开源与专有多模态模型,发现显著局限:即使最先进的模型如Gemini 2.5 Flash和Audio Flamingo 3,准确率也仅为59.2%和51.7%,在多个类别中接近随机表现。深入分析揭示了具体缺陷并提供新洞见,为社区提升未来AI系统向音频通用智能演进提供可操作建议。基准与代码已公开于https://sonalkum.github.io/mmau-pro。
原文摘要 · Abstract (English)
Audio comprehension-including speech, non-speech sounds, and music-is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio understanding to qualify as generally intelligent. However, evaluating auditory intelligence comprehensively remains challenging. To address this gap, we introduce MMAU-Pro, the most comprehensive and rigorously curated benchmark for assessing audio intelligence in AI systems. MMAU-Pro contains 5,305 instances, where each instance has one or more audios paired with human expert-generated question-answer pairs, spanning speech, sound, music, and their combinations. Unlike existing benchmarks, MMAU-Pro evaluates auditory intelligence across 49 unique skills and multiple complex dimensions, including long-form audio comprehension, spatial audio reasoning, multi-audio understanding, among others. All questions are meticulously designed to require deliberate multi-hop reasoning, including both multiple-choice and open-ended response formats. Importantly, audio data is sourced directly ``from the wild" rather than from existing datasets with known distributions. We evaluate 22 leading open-source and proprietary multimodal AI models, revealing significant limitations: even state-of-the-art models such as Gemini 2.5 Flash and Audio Flamingo 3 achieve only 59.2% and 51.7% accuracy, respectively, approaching random performance in multiple categories. Our extensive analysis highlights specific shortcomings and provides novel insights, offering actionable perspectives for the community to enhance future AI systems' progression toward audio general intelligence. The benchmark and code is available at https://sonalkum.github.io/mmau-pro.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。