arXiv:2510.20867cs.LGcs.AI2025-10被引 12

通过过程奖励提升音频大模型推理能力,解决越想越错的问题。

Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process Rewards

  • 用多维度奖励机制直接优化推理过程,而非只看结果
  • 在MMAU Test-mini上超越Gemini 2.5 Pro和GPT-4o Audio,接近人类水平
  • 发现模型专属推理“黄金区间”,长链推理反而更优

音频大语言模型中的推理作用长期被忽视,引入推理过程常导致性能下降,即测试时反向缩放现象:推理链越长,结果越差。我们证明这并非推理本质缺陷,而是训练不足所致——缺乏引导的推理易产生幻觉、不一致,错误累积。为此提出CESAR(一致、有效、可扩展的音频推理者),从结果验证转向推理过程奖励。其在线强化学习框架采用组相对策略优化,结合多维奖励,激励正确性、格式、一致性、结构化分析模式、因果推理、领域知识融合及合理推理深度。CESAR解决测试时反向缩放,使推理从负累转为增益,揭示各模型独有的“推理甜点区”,性能随推理长度提升而达到峰值。在MMAU Test-mini上显著优于Gemini 2.5 Pro和GPT-4o Audio,MMSU推理任务接近人类水平。通过AI裁判评估与定性对比,双重验证推理质量提升。更重要的是,增强推理同时促进多模态推理与感知能力协同进化。总体而言,CESAR建立了一套可扩展的音频大模型鲁棒推理开发范式。

原文摘要 · Abstract (English)

The role of reasoning in Audio Large Language Models remains widely underexplored, as introducing a reasoning process often degrades rather than improves performance during inference, a phenomenon we term test-time inverse scaling, where longer reasoning chains yield progressively worse results. We demonstrate that this stems not from fundamental limitations of reasoning itself, but from inadequate training: models without proper guidance for the reasoning process produce hallucinatory, inconsistent reasoning that accumulates errors over longer chains. To address these challenges, we introduce CESAR (Consistent, Effective, and Scalable Audio Reasoners), shifting from outcome verification to rewarding the reasoning process. Our online reinforcement learning framework employs Group Relative Policy Optimization with a multi-faceted reward suite that incentivizes not only correctness and format but also consistency, structured analytical patterns, causal reasoning, domain-knowledge integration, and calibrated reasoning depth. CESAR resolves test-time inverse scaling, transforming reasoning from detriments into gains while revealing model-specific ``reasoning sweet spots", where performance peaks during test-time scaling. We achieve state-of-the-art results on MMAU Test-mini, substantially outperforming Gemini 2.5 Pro and GPT-4o Audio, and near-human-level performance on MMSU reasoning tasks. Through AI-as-judge evaluations and qualitative comparisons, we provide both quantitative and qualitative validation of our improved reasoning quality. Importantly, enhanced reasoning creates synergistic effects, simultaneously improving multimodal reasoning and perception capabilities. Overall, CESAR establishes a principled method for developing robust and scalable reasoning in Audio LLMs.

音频LLM推理增强强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。