首个评估音频推理过程质量的挑战赛,推动可解释音频智能发展
The Interspeech 2026 Audio Reasoning Challenge: Evaluating Reasoning Process Quality for Audio Reasoning Models and Agents
- 设计基于事实性与逻辑性的实例级评测协议
- 多智能体系统在推理质量上领先,单模型通过强化学习快速进步
- 适合关注可解释音频模型与多模态推理的研究者
近期大型音频语言模型在理解能力上表现优异,但普遍缺乏透明的推理过程。为解决这一‘黑箱’问题,我们在2026年国际语音通信大会(Interspeech 2026)组织了首个专注于音频领域链式思维(Chain-of-Thought, CoT)质量评估的共享任务。挑战赛引入了MMAR-Rubrics——一种新的实例级评测协议,用于评估推理链的事实性和逻辑性。比赛包含单模型与智能体两个赛道,吸引了来自18个国家和地区的156支队伍参与。结果表明,当前智能体系统在推理质量上领先,其优势源于迭代工具调度与跨模态分析;同时,单模型正通过强化学习和复杂数据流水线快速提升。本文详细介绍了挑战赛的设计、方法,并对前沿系统进行了全面分析,为可解释音频智能提供了新洞见。
原文摘要 · Abstract (English)
Recent Large Audio Language Models (LALMs) excel in understanding but often lack transparent reasoning. To address this "black-box" limitation, we organized the Audio Reasoning Challenge at Interspeech 2026, the first shared task dedicated to evaluating Chain-of-Thought (CoT) quality in the audio domain. The challenge introduced MMAR-Rubrics, a novel instance-level protocol assessing the factuality and logic of reasoning chains. Featured Single Model and Agent tracks, the competition attracting 156 teams from 18 countries and regions. Results show agent systems currently lead in reasoning quality, utilizing iterative tool orchestration and cross-modal analysis. Besides, single models are rapidly advancing via reinforcement learning and sophisticated data pipeline. We details the challenge design, methodology, and a comprehensive analysis of state-of-the-art systems, providing new insights for explainable audio intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。