arXiv:2602.13685cs.SDcs.AI2026-02被引 4

让音频模型学会何时调用外部工具,提升复杂听觉推理能力。

AuTAgent: A Reinforcement Learning Framework for Tool-Augmented Audio Reasoning

  • 用强化学习决定何时调用外部工具,避免信息过载。
  • 在MMAU和MMAR上分别提升4.20%~9.80%准确率。
  • 适合需要精准音频分析的智能系统研发者。

大型音频语言模型(LALMs)在感知任务中表现优异,但在需精确声学测量的复杂推理任务中存在瓶颈。虽然外部工具可提取如节拍、音高这类细粒度特征,但有效整合仍具挑战:盲目使用所有工具会导致信息过载,而基于提示的选择无法评估工具在不同上下文中的实际价值。为此,我们提出AuTAgent(音频工具代理),一种基于强化学习的框架,自动学习何时以及调用何种工具。通过采用稀疏反馈训练策略与新颖的差分奖励机制,该代理能过滤无关工具,并仅在外部辅助带来净性能提升时才调用。实验表明,AuTAgent有效弥补了LALMs的表征局限,提供可验证的声学证据。在开源与闭源模型上,于MMAU Test-mini和MMAR基准测试中分别实现4.20%/6.20%和9.80%/8.00%的准确率提升。此外,迁移性实验显示其卓越泛化能力。本文强调外部工具在增强音频模型推理中的互补作用。

原文摘要 · Abstract (English)

Large Audio Language Models (LALMs) excel at perception but struggle with complex reasoning requiring precise acoustic measurements. While external tools can extract fine-grained features like exact tempo or pitch, effective integration remains challenging: naively using all tools causes information overload, while prompt-based selection fails to assess context-dependent utility. To address this, we propose AuTAgent (Audio Tool Agent), a reinforcement learning framework that learns when and which tools to invoke. By employing a sparse-feedback training strategy with a novel Differential Reward mechanism, the agent learns to filter out irrelevant tools and invokes external assistance only when it yields a net performance gain over the base model. Experimental results confirm that AuTAgent complements the representation bottleneck of LALMs by providing verifiable acoustic evidence. It improves accuracy by 4.20% / 6.20% and 9.80% / 8.00% for open-source and closed-source backbones on the MMAU Test-mini and the MMAR benchmarks, respectively. In addition, further experiments demonstrate exceptional transferability. We highlight the complementary role of external tools in augmenting audio model reasoning.

音频推理强化学习工具增强多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。