通过置信度感知强化学习,提升大模型推理能力
ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning

- 用词元置信度信号改进强化学习奖励机制
- 在多个模型规模下平均提升2.3%-4.0%
- 适合关注模型推理可靠性与稳定性研究者
强化学习从可验证奖励(RLVR)是提升大语言模型(LLM)推理能力的关键范式,但受限于稀疏的二值奖励及对模型内部不确定性的忽视。本文提出ConSteer-RL框架,将基于模型对数概率的词元级置信度信号融入RLVR训练。具体地,在组相对策略优化(GRPO)基础上,通过聚合词元概率生成标量置信度得分,并引入意识增强型奖励塑造机制,惩罚过度自信的错误,同时强化正确且自信的推理。实验表明,ConSteer-RL在不同模型规模下均显著优于强基线GRPO,平均提升2.3%至4.0%。
原文摘要 · Abstract (English)
Reinforcement Learning from Verifiable Rewards (RLVR) has recently become a key paradigm for improving the reasoning abilities of Large Language Models (LLMs), yet it remains limited by sparse binary rewards and its ignorance of model-internal uncertainty. In this paper, we propose ConSteer-RL, a simple yet effective framework that integrates token-level confidence signals derived from model log-probabilities into RLVR training. Specifically, building upon the Group Relative Policy Optimization (GRPO) framework, we construct a confidence-aware reward by aggregating per-token probabilities into a scalar confidence score and incorporating it into an awareness-based reward shaping mechanism that penalizes overconfident errors while reinforcing correct and confident reasoning. Experimental results demonstrate that ConSteer-RL consistently outperforms strong GRPO baselines, achieving average improvements of 2.3%-4.0% across different model scales.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。