arXiv:2508.11800cs.LGcs.AI2025-08被引 21

GRPO让语言模型对随机结果过度自信,改方法可修复校准问题。

Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes

  • 用组内标准归一化导致模型预测过度自信
  • 移除该归一化后,GRPO预测校准度显著提升
  • 适合关注模型可信度与随机推理的AI研究者

强化学习(RL)在数学等确定性任务中已证明能有效提升语言模型准确性。本文检验当前RL方法在具有随机结果的可验证领域(如生物实验)中的表现。通过合成数据和真实生物实验的应用,发现组相对策略优化(GRPO)会导致二元随机结果的概率预测过度自信,而近端策略优化(PPO)与留一法强化学习(RLOO)则生成校准良好的模型。研究显示,移除GRPO中的组标准归一化可解决其校准偏差,并提供了理论解释:该归一化是导致过度自信的原因。结果为避免在GRPO中使用标准归一化提供新证据,助力强化学习在非确定性推理任务中的应用。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has proven remarkably effective at improving the accuracy of language models in verifiable and deterministic domains like mathematics. Here, we examine if current RL methods are also effective at optimizing language models in verifiable domains with stochastic outcomes, like scientific experiments. Through applications to synthetic data and real-world biological experiments, we demonstrate that Group Relative Policy Optimization (GRPO) induces overconfident probability predictions for binary stochastic outcomes, while Proximal Policy Optimization (PPO) and REINFORCE Leave-One-Out (RLOO) yield well-calibrated models. We show that removing group standard normalization in GRPO fixes its miscalibration and provide a theoretical explanation for why normalization causes overconfidence. Our results provide new evidence against the use of standard normalization in GRPO and help pave the way for applications of RL for reasoning language models beyond deterministic domains.

强化学习模型校准语言模型随机推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。