arXiv:2605.28480eess.AScs.SD2026-05被引 2

让音频智能体只在必要时调用外部工具,提升判断可靠性。

Audio-Mind: An Auditable Agentic Framework for Audio Understanding

论文配图:Audio-Mind: An Auditable Agentic Framework for Audio Understanding
图 1 · 摘自论文原文
  • 根据初始证据充足度动态决定是否调用外部工具
  • 在两个数据集上分别达到80.4%和82.8%准确率
  • 生成可审计的推理过程,适合需要透明性的音频问答场景

音频智能体通过将音频问题分解为工具调用、中间证据和迭代推理步骤,扩展了大音频语言模型(LALMs)的能力。然而,随着LALMs性能增强,核心挑战已从实现工具使用转向判断何时获取代理证据能真正提升音频理解。本文提出Audio-Mind,一个可审计、可插拔的条件性证据获取框架。该框架动态结合强前端模型与规划引导的工具调用,在初始证据充分时保留前端整体判断,仅在存在未解决证据缺口时获取有界外部证据。在MMAR和MSU-Bench上的实验表明,Audio-Mind优于现有音频智能体基线,分别达到80.4%和82.8%的准确率。匹配骨干对比显示:在强音频前端下,若工作流不保留前端的整体音频感知判断,代理分解可能成为协调瓶颈。除精度外,Audio-Mind还能生成高质量、可审计的推理轨迹,揭示不确定性、工具证据与答案依据,为更可靠的音频问答标注与错误分析提供基础。

原文摘要 · Abstract (English)

Audio agents extend large audio-language models (LALMs) by decomposing audio questions into tool calls, intermediate evidence, and iterative reasoning steps. However, as LALMs become stronger, the key challenge shifts from enabling tool use to determining when agentic evidence acquisition genuinely benefits audio understanding. We propose Audio-Mind, an auditable and pluggable framework for conditional evidence acquisition in audio understanding. Audio-Mind dynamically combines a strong frontend with planner-guided tool use, preserving frontend judgment when initial evidence is sufficient while acquiring bounded external evidence for questions with unresolved evidence gaps. Experiments on MMAR and MSU-Bench show that Audio-Mind outperforms prior audio-agent baselines, reaching 80.4% accuracy on MMAR and 82.8% accuracy on MSU-Bench. A matched-backbone comparison highlights why this design matters: under strong audio frontends, agentic decomposition can become an orchestration bottleneck when the workflow does not preserve the frontend's holistic audio-grounded judgment. Beyond accuracy, Audio-Mind produces higher-quality, auditable reasoning traces that expose uncertainty, tool evidence, and answer rationales, offering a potential basis for more reliable audio-QA annotation and error analysis.

音频理解智能体框架可审计性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。