arXiv:2607.19824cs.AIcs.CL2026-07

让大模型推理过程更符合人类偏好,提升对齐效果。

Rewarding Better Thinking for LLM Preference Alignment

论文配图:Rewarding Better Thinking for LLM Preference Alignment
图 1 · 摘自论文原文
  • 用思维清单评估推理过程是否覆盖偏好关键点。
  • 在5个模型上显著提升对齐性能,尤其在复杂指令下。
  • 适合关注推理质量与人类价值观对齐的研究者。

大语言模型偏好对齐旨在使模型在多种用户指令下符合人类偏好。强化学习已成为主流后训练方法,但现有代理奖励多为结果层面,仅评估最终回复,对推理轨迹指导有限。这导致多个回复得分相近时,信用分配粗糙,轨迹级偏好难以明确。为此,我们提出思维检查表奖励(TCR),一种面向过程的强化学习对齐奖励。TCR将偏好对转化为特定样本的思维检查表,评估生成的推理轨迹是否涵盖偏好隐含的关键考虑。为减少与结果奖励的重叠,TCR进一步引入指数移动平均(EMA)残差形式,分离出超出结果可预测性的额外思维贡献。在三个模型家族的五个模型上实验表明,TCR在多个基准测试中一致提升对齐性能;消融实验验证了EMA残差和样本特异性检查表监督的重要性。

原文摘要 · Abstract (English)

LLM preference alignment aims to optimize models toward human preferences across diverse user instructions. Reinforcement learning has become a major post-training approach for this goal, but existing proxy rewards are often outcome-level, mainly evaluating the final response while providing limited guidance for the reasoning trajectory. This can make credit assignment coarse when multiple responses receive similar final scores, leaving trajectory-level preferences under-specified. To address this limitation, we propose Thinking Checklist Reward (TCR), a process-oriented reward for RL-based preference alignment. TCR converts preference pairs into sample-specific thinking checklists and uses them to evaluate whether the generated reasoning trace addresses the preference-implied considerations. To reduce overlap with outcome-level supervision, TCR further introduces an exponential moving average (EMA) residual formulation to isolate a complementary thinking surplus beyond what is predictable from the outcome reward. Experiments on five models from three model families show that TCR consistently improves alignment performance across diverse benchmarks, with ablations further validating the importance of EMA-based residual formulation and sample-specific checklist supervision.

偏好对齐强化学习推理评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。