arXiv:2508.08039cs.SDcs.CL2025-08被引 25

用强化学习让音频大模型学会何时、如何思考,提升听觉语言推理能力。

Audio-Thinker: Guiding Audio Language Model When and How to Think via Reinforcement Learning

  • 通过自适应思考精度奖励,动态调整复杂任务的推理策略。
  • 引入外部奖励模型,显著提升推理过程的一致性与质量。
  • 适合研究音频语言模型推理机制或需要强泛化能力的场景。

近期大型语言模型、多模态大语言模型及大型音频语言模型(LALMs)通过基于规则的奖励强化学习,显著提升了推理能力。然而,显式推理过程在音频问答任务中尚未带来明显优势,深度推理的有效利用仍是开放挑战,现有LALMs仍远未达到人类水平的听觉-语言推理能力。为解决此问题,我们提出Audio-Thinker,一种增强LALMs推理能力的强化学习框架,重点提升其适应性、一致性和有效性。该方法引入自适应思考精度奖励,使模型能根据任务复杂度动态调整推理策略;同时结合外部奖励模型评估整体推理一致性与质量,并辅以基于思考的奖励,帮助模型在训练中区分有效与错误推理路径。实验结果表明,Audio-Thinker在多个基准任务上优于现有推理导向的LALMs,展现出更优的推理与泛化能力。

原文摘要 · Abstract (English)

Recent advancements in large language models, multimodal large language models, and large audio language models (LALMs) have significantly improved their reasoning capabilities through reinforcement learning with rule-based rewards. However, the explicit reasoning process has yet to show significant benefits for audio question answering, and effectively leveraging deep reasoning remains an open challenge, with LALMs still falling short of human-level auditory-language reasoning. To address these limitations, we propose Audio-Thinker, a reinforcement learning framework designed to enhance the reasoning capabilities of LALMs, with a focus on improving adaptability, consistency, and effectiveness. Our approach introduces an adaptive think accuracy reward, enabling the model to adjust its reasoning strategies based on task complexity dynamically. Furthermore, we incorporate an external reward model to evaluate the overall consistency and quality of the reasoning process, complemented by think-based rewards that help the model distinguish between valid and flawed reasoning paths during training. Experimental results demonstrate that our Audio-Thinker model outperforms existing reasoning-oriented LALMs across various benchmark tasks, exhibiting superior reasoning and generalization capabilities.

音频推理强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。