arXiv:2604.18187cs.SDcs.CL2026-04被引 2

用强化学习让音频模型生成有声学依据的推理链,效果领先。

Audio-DeepThinker: Progressive Reasoning-Aware Reinforcement Learning for High-Quality Chain-of-Thought Emergence in Audio Language Models

论文配图:Audio-DeepThinker: Progressive Reasoning-Aware Reinforcement Learning for High-Quality Chain-of-Thought Emergence in Audio Language Models
图 1 · 摘自论文原文
  • 设计混合奖励机制,结合逻辑评估与语义相似性,精准优化推理质量。
  • 分两阶段渐进训练,纯强化学习使模型自发产生高质量推理链。
  • 在多个音频推理数据集上刷新纪录,适合研究可解释性与智能系统设计者。

大型音频语言模型在音频理解方面取得显著进展,但多为感知-回答系统,缺乏显式推理过程。现有增强音频推理的方法或依赖受监督的思维链微调(受限于数据质量),或使用粗粒度奖励的强化学习,无法直接评估推理质量。因此生成的推理链虽结构良好,却缺乏具体声学依据。本文提出 Audio-DeepThinker 框架,核心包含两点:一是引入混合推理相似性奖励,结合大模型评估逻辑路径一致性、关键步骤覆盖率与分析深度,以及嵌入相似性组件以保证与参考推理链的语义对齐;二是提出渐进式双阶段课程学习,仅通过纯强化学习探索,无需任何监督推理微调,从无思维链能力的指令微调模型中催生高质量思维链。第一阶段在基础音频问答任务上使用混合奖励训练基础推理模式,第二阶段转向声学挑战性边界案例,仅用大模型奖励提升推理多样性。Audio-DeepThinker 在 MMAR(74.0%)、MMAU-test-mini(78.5%)和 MMSU(77.26%)上达到当前最优,获 Interspeech 2026 音频推理挑战赛单模型赛道第一名。可解释性分析显示,强化学习主要重塑顶层 MoE 门控机制,推理标记逐步在上层 Transformer 层中凝练,揭示了通过探索实现音频推理涌现的机制。

原文摘要 · Abstract (English)

Large Audio-Language Models (LALMs) have made significant progress in audio understanding, yet they primarily operate as perception-and-answer systems without explicit reasoning processes. Existing methods for enhancing audio reasoning rely either on supervised chain-of-thought (CoT) fine-tuning, which is limited by training data quality, or on reinforcement learning (RL) with coarse rewards that do not directly evaluate reasoning quality. As a result, the generated reasoning chains often appear well-structured yet lack specific acoustic grounding. We propose Audio-DeepThinker, a framework built on two core ideas. First, we introduce a hybrid reasoning similarity reward that directly supervises the quality of generated reasoning chains by combining an LLM evaluator assessing logical path alignment, key step coverage, and analytical depth with an embedding similarity component enforcing semantic alignment with reference reasoning chains. Second, we propose a progressive two-stage curriculum that enables high-quality CoT reasoning to emerge through pure RL exploration, without any supervised reasoning fine-tuning, from an instruction-tuned model that possesses no prior chain-of-thought capability. Stage 1 trains on foundational audio QA with the hybrid reward to foster basic reasoning patterns, while Stage 2 shifts to acoustically challenging boundary cases with an LLM-only reward for greater reasoning diversity. Audio-DeepThinker achieves state-of-the-art results on MMAR (74.0%), MMAU-test-mini (78.5%), and MMSU (77.26%), winning 1st Place in the Interspeech 2026 Audio Reasoning Challenge (Single Model Track). Interpretability analyses further reveal that RL training primarily reshapes upper-layer MoE gating mechanisms and that reasoning tokens crystallize progressively in the upper transformer layers, offering mechanistic insights into how audio reasoning emerges through exploration.

音频推理强化学习思维链可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。