arXiv:2510.08899cs.LGcs.AI2025-10被引 3

提出新方法提升大模型推理中的步骤贡献评估精度

Pinpointing crucial steps: Attribution-based Credit Assignment for Verifiable Reinforcement Learning

  • 基于归因分析动态调节策略熵,缓解过早收敛问题
  • 在AIME、MATH等竞赛题上超越现有最优方法
  • 适合研究大模型复杂推理与可验证奖励机制的学者

尽管带有可验证奖励的强化学习(RLVR)提升了大语言模型的复杂推理能力,但现有方法难以平衡探索与利用,导致中间步骤信用分配不准确及熵过早坍塌,限制模型性能。为此,我们提出基于归因的策略优化贡献(ACPO),一个分阶段框架,融合难度感知课程设计。ACPO通过轨迹语义分割和基于归因的表示动态调节策略熵,增强探索;同时采用分解式奖励系统,精确量化每个推理步骤的层次化贡献,实现精准信用分配。在AIME、MATH和AMC等挑战性基准上的大量实验表明,ACPO显著优于现有最先进方法。

原文摘要 · Abstract (English)

While Reinforcement Learning with Verifiable Rewards (RLVR) enhances complex reasoning in LLMs, current methods struggle to balance exploration and exploitation. This leads to critical issues like inaccurate credit assignment for intermediate steps and premature entropy collapse, limiting model performance. To address this, we introduce Attribution-based Contribution to Policy Optimization (ACPO), a phased framework that incorporates a difficulty-aware curriculum. ACPO improves exploration by using trajectory semantic segmentation and an attribution-based representation to dynamically regulate policy entropy, thus mitigating its collapse. Concurrently, it enhances exploitation with a factorized reward system that precisely quantifies the hierarchical contribution of each reasoning step, ensuring accurate credit assignment. Extensive experiments on challenging benchmarks, including AIME, MATH, and AMC, demonstrate that ACPO significantly outperforms existing state-of-the-art approaches.

强化学习大模型推理信用分配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。