arXiv:2511.21005cs.AIcs.IR2025-11被引 3

用模型自评概率优化强化学习,提升大模型推理稳定性。

ICPO: Intrinsic Confidence-Driven Group Relative Preference Optimization for Efficient Reinforcement Learning

  • 基于响应生成概率差计算偏好优势分,动态引导探索。
  • 在4个通用和3个数学基准上优于GRPO,推理能力显著提升。
  • 适合需要稳定高阶推理的RLHF场景,尤其关注过拟合问题。

基于可验证奖励的强化学习(RLVR)在提升大语言模型(LLM)推理能力方面展现出巨大潜力。然而,现有方法常受粗粒度奖励、奖励噪声及低效探索等问题制约,导致训练不稳定和熵崩溃。为此,我们提出内在置信度驱动的组相对偏好优化方法(ICPO)。其核心思想是:模型生成不同响应的概率能直接反映其对推理过程的自我评估。受偏好建模启发,ICPO通过比较同一输入下多个响应的生成概率差异,计算出偏好优势分数,并将其与可验证奖励结合,指导探索过程。实验发现,该分数不仅能缓解奖励粗糙和噪声问题,还能有效抑制过度自信错误,增强被低估的高质量响应的相对优势,防止模型过度拟合特定策略。在四个通用领域基准和三个数学基准上的全面实验表明,ICPO相比GRPO持续提升推理性能。

原文摘要 · Abstract (English)

Reinforcement Learning with Verifiable Rewards (RLVR) demonstrates significant potential in enhancing the reasoning capabilities of Large Language Models (LLMs). However, existing RLVR methods are often constrained by issues such as coarse-grained rewards, reward noise, and inefficient exploration, which lead to unstable training and entropy collapse. To address this challenge, we propose the Intrinsic Confidence-Driven Group Relative Preference Optimization method (ICPO). The intuition behind it lies in the fact that the probabilities of an LLM generating different responses can inherently and directly reflect its self-assessment of the reasoning process. Inspired by the idea of preference modeling, ICPO calculates a preference advantage score for each response by comparing the relative generation probabilities of multiple responses under the same input prompt, and integrates this score with verifiable rewards to guide the exploration process. We have discovered that the preference advantage score not only alleviates the issues of coarse-grained rewards and reward noise but also effectively curbs overconfident errors, enhances the relative superiority of undervalued high-quality responses, and prevents the model from overfitting to specific strategies. Comprehensive experiments across four general-domain benchmarks and three mathematical benchmarks demonstrate that ICPO steadily boosts reasoning compared to GRPO.

强化学习大模型推理偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。