arXiv:2606.05932cs.AIcs.LG2026-06

拆解强化学习中奖励设计与自我一致性激发的混淆效应,揭示常见评估指标的系统偏差。

A Pre-Registered Causal Partition of Self-Consistency Elicitation and Reward Design in RLVR

  • 提出因果分解框架,分离奖励设计与自我一致性激发的影响
  • 实测显示奖励设计贡献仅占0.05至0.139,自我一致性项可反号
  • 提供可复现工具,适用于对齐研究的审计验证

强化学习从可验证奖励(RLVR)在奖励信号虚假时仍能提升推理能力——即给予群体多数答案信用而非真实验证结果。从业者常将 naive = acc(TRUE) - acc(RANDOM) 视为奖励设计效应。我们证明该估计量存在系统性偏差:它混淆了自我一致性激发(通过多数伪奖励使策略趋向其众数答案)与真正的奖励设计信号。利用受控的表格型GRPO模拟器,我们推导出精确的递推分解:total = null + elicit + rd,并在五个先验强度水平下测量各项。奖励设计部分占原估计量的比例从弱先验(ps=0.20)下的0.139降至强先验(ps=0.80)下的0.05,而激发项在自我一致性交叉点发生符号反转。预注册的2x2x2因子实验确认非加性关系(交互比0.385;AxC效应-0.089)。点-界试点门分析表明,强先验区域为点识别,接近交叉区域仅能边界识别。对两篇已发表研究的重新审计分别得出‘激发主导’(激发占比0.98)和‘奖励设计主导’(奖励设计占比1.18)结论,验证了该分解方法的诊断价值。我们预先承诺无论结果如何均提交;无翻转亦为有效发现。发布一键式可复用工具,供任意对齐论文运行相同审计。

原文摘要 · Abstract (English)

Reinforcement learning from verifiable rewards (RLVR) improves reasoning even when the reward signal is spurious -- assigning credit to the group-plurality answer rather than a ground-truth verifier. Practitioners commonly interpret naive = acc(TRUE) - acc(RANDOM) as the reward-design effect. We prove this estimand is systematically biased: it conflates self-consistency elicitation (sharpening the policy toward its modal answer via majority pseudo-reward) with genuine reward-design signal. Using a controlled tabular-GRPO simulator we derive an exact telescoping decomposition total = null + elicit + rd and measure each term across five prior-strength levels. The reward-design fraction of the naive estimator ranges from 0.139 at weak prior (ps=0.20) to 0.05 at strong prior (ps=0.80), with the elicitation term flipping sign at the self-consistency crossover. A pre-registered 2x2x2 factorial confirms non-additivity (interaction ratio 0.385; AxC effect -0.089). A points-vs-bounds pilot gate shows strong-prior regimes are point-identified while near-crossover regimes are only bounded. Re-audits of two named published results yield ELICITATION DOMINATED (elicitation share 0.98) and REWARD DESIGN DOMINATED (rd share 1.18) verdicts respectively, demonstrating the diagnostic value of the partition. We pre-commit to submit regardless of flip outcome; a non-flip is a finding of equal standing. We release a reusable one-command harness for any alignment paper to run the same audit.

强化学习因果分解模型对齐可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。