arXiv:2606.15141eess.AScs.AI2026-06中稿 · Interspeech 2026被引 1

让语音问答有逻辑可查:分步推理并自我验证。

EChO-Agent: Evidence Chain Orchestration Agent for Audio Reasoning

论文配图:EChO-Agent: Evidence Chain Orchestration Agent for Audio Reasoning
图 1 · 摘自论文原文
  • 将复杂语音问答拆解为规划、工具执行、证据整合和答案验证流程。
  • 在MMAR基准上准确率与评分均优于基线,证据整合是关键提升点。
  • 适合需要可解释性推理的语音理解任务,如智能助手、听障辅助。

尽管语言音频模型(LALMs)在语音问答中展现出潜力,但在处理复杂音频推理时,往往无法聚焦问题相关音频段落,且缺乏清晰、可核查的推理过程。强化学习与工具增强提示虽有助于模型更好地关联问题与音频,但缺少可靠方法来理解、整合并自验证音频片段。为此,我们提出EChO-Agent,一个模块化代理框架,将复杂语音问答重构为规划、工具执行、证据整合与答案验证的工作流。在MMAR基准上的实验表明,EChO-Agent在准确率和评分上均优于基线;消融研究进一步显示,证据整合是核心提升因素。

原文摘要 · Abstract (English)

While LALMs show promise on audio question answering, they fail to focus on question-relevant segments of audio and provide a clear, checkable reasoning process when dealing with complex audio reasoning. Reinforcement learning and tool-augmented prompting can help models better relate questions to audio but lack a reliable way to understand, integrate, and self-verify audio segments. To address this gap, we present EChO-Agent, a modular agent framework that reformulates complex audio QA as a planning, tool execution, evidence integration, and answer verification workflow. Experiments on MMAR benchmark show EChO-Agent improves both accuracy and rubric scores over baseline and ablation studies show evidence integration is the key factor.

语音推理可解释性智能代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。