arXiv:2606.09124cs.AI2026-06

用后悔最小化重构人类反馈,让大模型更懂真实偏好

A Regret Minimization Framework on Preference Learning in Large Language Models

论文配图:A Regret Minimization Framework on Preference Learning in Large Language Models
图 1 · 摘自论文原文
  • 将人类反馈视为对行为次优性的相对评估,而非直接奖励
  • 在数学推理和人类偏好数据集上表现优于传统方法
  • 适合追求模型与人类真实意图对齐的研究者

基于可验证奖励的强化学习(RLVR)通过任务特定的验证器提供自动化正确性信号,推动了复杂推理任务的发展。然而,许多现实语言任务难以配备可靠的验证器,促使研究转向从人类反馈中进行强化学习(RLHF)。本文指出,需重新审视人类反馈的含义。为此提出基于后悔最小化的偏好优化方法(RePO),将RLHF重新定义为后悔最小化而非奖励最大化。人类偏好往往源于对未来结果的预期及对替代行为的反事实比较,而非即时、与结果无关的效用。RePO通过建模行为条件下的相对次优性来捕捉这一机制。在数学推理基准和人类偏好数据集上的实验表明,RePO持续取得性能提升,证明其是一种有效且符合人类意图的大语言模型训练方法。

原文摘要 · Abstract (English)

Reinforcement learning with verifiable rewards (RLVR) has enabled progress on reasoning-intensive tasks by relying on task-specific verifiers that provide automated correctness signals. However, many realistic language tasks are difficult to equip with reliable verifiers, motivating a growing reliance on reinforcement learning from human feedback (RLHF). In this setting, we argue that a closer examination of how human feedback should be interpreted is essential. We introduce Regret-based Preference Optimization $(\textbf{RePO})$, which reframes RLHF through $\textit{regret minimization}$ rather than reward maximization. Human preferences are often shaped by $\textit{prospective}$ anticipation of outcomes and $\textit{counterfactual}$ comparisons to alternative behaviors, rather than by immediate, outcome-independent utility. $\textbf{RePO}$ captures this structure by modeling preferences as behavior-conditioned assessments of relative suboptimality. Experiments on mathematical reasoning benchmarks and human preference datasets demonstrate consistent performance gains, indicating that $\textbf{RePO}$ is an effective and human-aligned approach for training large language models.

偏好学习强化学习人类对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。