让大模型学会用外部工具听音,省下600倍数据
AudioRouter: Data Efficient Audio Understanding via RL based Dual Reasoning
- 用强化学习决定何时调用外部音频工具
- 仅需常规方法1/600数据就达成更好性能
- 适合想高效训练语音理解模型的研究者
大型音频语言模型(LALMs)在音频理解与推理方面展现出强大能力,但在细粒度听觉感知上表现仍不可靠,现有方法大多依赖海量数据训练以内化感知能力。本文提出AudioRouter,一种基于强化学习的框架,使LALMs通过学习何时以及如何使用外部音频工具来提升音频理解能力。不同于将工具使用与推理紧密耦合,AudioRouter将工具使用建模为显式的决策问题,并在保持底层推理模型冻结的前提下优化轻量级路由策略。实验结果表明,AudioRouter在标准音频理解基准上实现显著提升,且学习工具使用所需训练数据最多减少600倍。这些发现表明,学习有效工具使用为内化感知能力提供了一种数据高效且可扩展的替代方案。
原文摘要 · Abstract (English)
Large Audio Language Models (LALMs) have demonstrated strong capabilities in audio understanding and reasoning. However, their performance on fine grained auditory perception remains unreliable, and existing approaches largely rely on data intensive training to internalize perceptual abilities. We propose AudioRouter, a reinforcement learning framework that enables LALMs to improve audio understanding by learning when and how to use external audio tools. Rather than tightly coupling tool usage with audio reasoning, AudioRouter formulates tool use as an explicit decision making problem and optimizes a lightweight routing policy while keeping the underlying reasoning model frozen. Experimental results show that AudioRouter achieves substantial improvements on standard audio understanding benchmarks while requiring up to 600x less training data to learn tool usage compared with conventional training paradigms. These findings suggest that learning effective tool usage offers a data efficient and scalable alternative to internalizing perceptual abilities in LALMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。