让语音模型学会调用工具、多轮推理,解决复杂音频任务。
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models

- 通过轨迹监督微调和强化学习,让模型掌握工具调用与交互策略。
- 在标准测试集上达到84.17分,在分布外任务上仍保持70.94分。
- 适合需要动态工具协作的智能语音助手、语音分析系统开发者。
复杂的声学问题可能需要模型执行声学操作、与外部工具交互,并基于生成的文本或处理后的音频结果进行推理,而非直接从固定音频输入中回答。我们研究此类问题为工具交互式音频推理,并开发了SpeechAgent-R,一个将内在多模态理解与外部技能和工具协调的音频代理。为支持该能力,我们构建了HIU-Corpus,包含65,492条交互轨迹和507.6小时音频,覆盖24个任务、8种技能和9种工具。SpeechAgent-R首先通过基于轨迹的监督微调学习结构化交互行为,再通过多轮强化学习优化决策。我们还引入HIU-Bench,用于联合评估任务性能、交互质量及对多样化任务设置的泛化能力。该基准包含1,395个样本,覆盖56个任务,包含分布内(ID)和分布外(OOD)划分,工具使用和工作流组合存在显著差异。SpeechAgent-R在ID任务上达到84.17分,在OOD任务上达到70.94分,相比同一代理框架下的基础模型分别提升15.40和14.23分。结果表明,学习技能与工具协调能显著提升音频代理应对多样任务场景和自适应工具交互的能力。
原文摘要 · Abstract (English)
Complex acoustic problems may require models to perform acoustic operations, interact with external tools and reason over the resulting textual or processed-audio observations rather than answer directly from a fixed audio input. We study such problems as tool-interactive audio reasoning and develop SpeechAgent-R, an audio agent that coordinates its intrinsic multimodal understanding with external skills and tools. To support this capability, we construct HIU-Corpus, comprising 65,492 interaction trajectories and 507.6 hours of audio across 24 tasks, 8 skills and 9 tools. SpeechAgent-R first learns structured interaction behaviors through trajectory-based supervised fine-tuning and then improves its decisions through multi-turn reinforcement learning. We further introduce HIU-Bench to jointly evaluate task performance, interaction quality and generalization to diverse task settings. It contains 1,395 samples across 56 tasks, including in-distribution (ID) and out-of-distribution (OOD) splits with substantial shifts in tool usage and workflow composition. SpeechAgent-R achieves 84.17 on ID tasks and 70.94 on OOD tasks, improving over the base model under the same agent harness by 15.40 and 14.23 points. These results demonstrate that learning skill and tool coordination improves audio agents' ability to handle diverse task settings and adaptive tool interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。