中文多场景动态音频推理新基准,挑战模型复杂听觉理解能力
CMDAR: A Chinese Multi-scene Dynamic Audio Reasoning Benchmark with Diverse Challenges
- 构建涵盖5类复杂推理的3000个问答对,模拟多源音频动态交互
- 主流模型在多音频多选题上准确率仅68%-77%,暴露出推理短板
- 适合研究多模态智能、语音理解与中文AI评测的学者参考
音频推理能力对AI代理在真实世界中的有效交互至关重要。现有基准多聚焦静态或单一场景下的英文音频数据,未能充分覆盖多说话人、事件演进和异构音频源交织的复杂情境。为此,我们提出CMDAR,一个面向中文场景的多场景动态音频推理评测基准。该基准包含3000个精心设计的问答对,关联多样音频片段,覆盖五类复杂推理任务,涵盖三种题型。我们在CMDAR上评估了26个先进音频语言模型,发现其在复杂推理任务中表现受限:Qwen2.5-Omni在CMDAR-main上达76.67%准确率,GPT-4o Audio为68.47%。但GPT-4o Audio在更难的多音频多选题和开放性任务上显著优于前者。我们进一步提供详细分析与未来模型发展的建议。
原文摘要 · Abstract (English)
The ability to reason from audio, including speech, environmental sounds, and music, is essential for AI agents to interact effectively in real-world scenarios. Existing benchmarks mainly focus on static or single-scene settings and English audio data and do not fully capture scenarios where multiple speakers, unfolding events, and heterogeneous audio sources interact. To address these challenges, we introduce CMDAR, a Chinese benchmark for evaluating models on complex, multi-scene, and dynamically evolving audio reasoning tasks. CMDAR comprises 3,000 carefully curated question-answer pairs linked to diverse audio clips, covering five categories of complex reasoning and spanning three question types. We benchmark 26 state-of-the-art audio language models on CMDAR and observe that they exhibit limitations in complex reasoning tasks. In CMDAR-main, Qwen2.5-Omni achieves 76.67% accuracy, whereas GPT-4o Audio reaches 68.47%. However, GPT-4o Audio substantially outperforms Qwen2.5-Omni on the more challenging multiple-choice with multiple audios and open-ended tasks. And we provide detail analysis corresponding suggestions for the future development of large audio language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。