arXiv:2510.22255cs.AIcs.CL2025-10被引 5

用模型自身信心变化做密集奖励,加速大模型推理训练。

PACR: Progressively Ascending Confidence Reward for LLM Reasoning

  • 基于模型对正确答案的信心上升趋势设计内在奖励
  • 减少探索轨迹数,更快达到奖励饱和,提升多任务表现
  • 适合需要高效推理训练的RLVR研究者

强化学习结合可验证奖励(RLVR)显著提升了大模型的推理能力,但其稀疏的、基于最终结果的奖励无法指导中间步骤,导致探索效率低下。本文提出渐进式上升信心奖励(PACR),一种直接从模型对正确答案信念演化中计算出的密集型内在奖励。PACR编码了这样一种归纳偏置:在合理的推理路径上,真实答案的概率应呈现总体上升趋势。我们通过实证与理论分析验证,这种归纳偏置能将探索搜索空间限制在逻辑更严谨的区域。实验表明,PACR能加速探索过程,在更少的推理轨迹下达到奖励饱和,并在多个基准测试中取得性能提升。结果表明,密集且模型内生的引导信号可使RLVR训练更高效、更可靠。

原文摘要 · Abstract (English)

Reinforcement Learning with Verifiable Rewards (RLVR) has significantly improved LLM reasoning, but its sparse, outcome-based reward provides no guidance for intermediate steps, slowing exploration. We propose Progressively Ascending Confidence Reward (PACR), a dense, model-intrinsic reward computed directly from the model's evolving belief in the correct answer. PACR encodes the inductive bias that, along a well-formed reasoning trajectory, the probability of the ground-truth answer should have a generally ascending trend. We provide empirical and theoretical analysis validating that such an inductive bias constrains the exploration search space to regions richer in logically sound reasoning. We demonstrate that PACR accelerates exploration, reaches reward saturation with fewer trajectories, and yields improvements on multiple benchmarks. Our results suggest that dense, model-intrinsic shaping signals can make RLVR training more effective and reliable.

大模型推理强化学习奖励设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。