arXiv:2512.16912cs.LGcs.AI2025-12中稿 · ICLR被引 25

用伪奖励提升大模型推理,揭示其背后的熵与剪裁机制。

Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward

  • 通过剪裁偏差降低策略熵,使模型输出更确定
  • 伪奖励可提升性能,而单纯熵最小化无效
  • 适合研究大模型强化学习与推理优化的读者

本文研究基于可验证奖励的强化学习(RLVR)中的探索-利用权衡问题,该框架旨在提升大语言模型(LLMs)的推理能力。近期研究表明,尽管看似矛盾,伪奖励(即奖励与真实答案无关)和熵最小化均能增强推理表现:前者抑制利用,后者抑制探索。本文聚焦两个核心问题:(i) 策略熵与性能的关系;(ii) 伪奖励是否真正带来收益,可能通过剪裁偏差与模型污染的相互作用。结果表明,在伪奖励下,剪裁偏差会降低策略熵,促使模型输出更自信、更确定,而单独的熵最小化无法实现性能提升。进一步提出奖励错配模型,解释为何伪奖励能在非污染场景下仍有效。研究厘清了伪奖励的增益机制,为更有效的RLVR训练提供了指导原则。

原文摘要 · Abstract (English)

This paper examines the exploration-exploitation trade-off in reinforcement learning with verifiable rewards (RLVR), a framework for improving the reasoning of Large Language Models (LLMs). Recent studies suggest that RLVR can elicit strong mathematical reasoning in LLMs through two seemingly paradoxical mechanisms: spurious rewards, which suppress exploitation by rewarding outcomes unrelated to the ground truth, and entropy minimization, which suppresses exploration by pushing the model toward more confident and deterministic outputs, highlighting a puzzling dynamic: both discouraging exploitation and discouraging exploration improve reasoning performance, yet the underlying principles that reconcile these effects remain poorly understood. We focus on two fundamental questions: (i) how policy entropy relates to performance, and (ii) whether spurious rewards yield gains, potentially through the interplay of clipping bias and model contamination. Our results show that clipping bias under spurious rewards reduces policy entropy, leading to more confident and deterministic outputs, while entropy minimization alone is insufficient for improvement. We further propose a reward-misalignment model explaining why spurious rewards can enhance performance beyond contaminated settings. Our findings clarify the mechanisms behind spurious-reward benefits and provide principles for more effective RLVR training.

强化学习大模型推理奖励设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。