构建首个覆盖多任务的音频理解与推理评测基准,推动模型向专家级能力迈进。
MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
- 设计10,000个精心筛选的音视频样本,涵盖语音、环境声与音乐
- 引入27项跨域技能测试,最大模型准确率仅52.97%
- 适合追求复杂音频推理能力的开发者与研究者
音频理解(包括语音、非语音声音和音乐)对人工智能体有效交互至关重要。我们提出MMAU,一个新型基准,用于评估多模态音频理解模型在需要专家级知识与复杂推理的任务上的表现。MMAU包含10,000个经过精心筛选的音频片段,配有人类标注的自然语言问答,覆盖语音、环境声和音乐。任务涵盖信息提取与推理,要求模型展现27种不同技能,涉及独特且具有挑战性的任务。与现有基准不同,MMAU强调领域特定知识下的高级感知与推理,使模型面临类似专家的实际挑战。我们评估了18个开源与专有(大型)音频-语言模型,揭示了MMAU带来的巨大挑战。值得注意的是,即使是最先进的Gemini Pro v1.5也仅达到52.97%的准确率,而最先进的开源模型Qwen2-Audio也仅为52.50%,表明仍有巨大提升空间。我们相信MMAU将推动音频与多模态研究社区开发出更先进的音频理解模型,以应对复杂音频任务。
原文摘要 · Abstract (English)
The ability to comprehend audio--which includes speech, non-speech sounds, and music--is crucial for AI agents to interact effectively with the world. We present MMAU, a novel benchmark designed to evaluate multimodal audio understanding models on tasks requiring expert-level knowledge and complex reasoning. MMAU comprises 10k carefully curated audio clips paired with human-annotated natural language questions and answers spanning speech, environmental sounds, and music. It includes information extraction and reasoning questions, requiring models to demonstrate 27 distinct skills across unique and challenging tasks. Unlike existing benchmarks, MMAU emphasizes advanced perception and reasoning with domain-specific knowledge, challenging models to tackle tasks akin to those faced by experts. We assess 18 open-source and proprietary (Large) Audio-Language Models, demonstrating the significant challenges posed by MMAU. Notably, even the most advanced Gemini Pro v1.5 achieves only 52.97% accuracy, and the state-of-the-art open-source Qwen2-Audio achieves only 52.50%, highlighting considerable room for improvement. We believe MMAU will drive the audio and multimodal research community to develop more advanced audio understanding models capable of solving complex audio tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。