通过自适应融合推理奖励,提升大模型用工具的能力
AWPO: Enhancing Tool-Use of Large Language Models through Adaptive Integration of Reasoning Rewards
- 用动态加权机制融合推理过程质量与结果奖励
- 40亿参数模型在多轮任务中超越Grok-4 16%准确率
- 适合需要高效、可靠工具调用的AI系统研发者
虽然强化学习在利用可验证结果奖励训练具备工具使用能力的大语言模型方面展现出潜力,但现有方法大多忽视了基于思维链质量的推理奖励对提升工具利用效果的可能。此外,简单地结合推理奖励与结果奖励可能导致性能下降或与主要优化目标冲突。为此,我们提出优势加权策略优化(AWPO),一种原则性强化学习框架,通过自适应地将推理奖励融入优势估计来改进工具使用表现。AWPO引入方差感知门控和难度感知加权,根据组内相对统计信息动态调节推理信号带来的优势,并采用定制化截断机制实现稳定优化。大量实验表明,AWPO在标准工具使用基准上达到当前最优性能,显著优于多个强基线模型,在复杂多轮场景中领先闭源模型。值得注意的是,凭借出色的参数效率,我们的40亿参数模型在多轮任务中的准确率比Grok-4高出16.0%,同时在分布外的MMLU-Pro基准上保持了良好的泛化能力。
原文摘要 · Abstract (English)
While Reinforcement Learning (RL) shows promise in training tool-use Large Language Models (LLMs) using verifiable outcome rewards, existing methods largely overlook the potential of reasoning rewards based on chain-of-thought quality for better tool utilization. Furthermore, naïvely combining reasoning and outcome rewards may yield suboptimal performance or conflict with the primary optimization objective. To address this, we propose Advantage-Weighted Policy Optimization (AWPO), a principled RL framework that adaptively integrates reasoning rewards into advantage estimation to improve tool-use performance. AWPO incorporates variance-aware gating and difficulty-aware weighting to adaptively modulate advantages from reasoning signals based on group-relative statistics, alongside a tailored clipping mechanism for stable optimization. Extensive experiments demonstrate that AWPO achieves state-of-the-art performance across standard tool-use benchmarks, significantly outperforming strong baselines and leading closed-source models in challenging multi-turn scenarios. Notably, with exceptional parameter efficiency, our 4B model surpasses Grok-4 by $16.0\%$ in multi-turn accuracy while preserving generalization capability on the out-of-distribution MMLU-Pro benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。