arXiv:2512.00709cs.AI2025-12AAAI被引 2

针对人类偏好翻转问题,提出更鲁棒的强化学习对齐算法。

When Human Preferences Flip: An Instance-Dependent Robust Loss for RLHF

  • 基于贝叶斯模型建模每条数据的翻转概率,实现实例级感知。
  • 在多个数据集上验证,相比基线方法显著提升对齐效果。
  • 适合需要高可靠性人类反馈训练的LLM应用场景。

大型语言模型对齐中数据质量至关重要。然而,在收集人类反馈时,偏好翻转现象普遍存在,导致标注数据污染;这要求对齐算法具备更强的抗翻转鲁棒性。为此,本文从强化学习与人类反馈(RLHF)视角出发,提出一种面向偏好翻转的翻转感知直接偏好优化(FA-DPO)算法。我们区分了人类意图建模与外部因素引发的偏好翻转两个阶段;在后一阶段,基于布拉德利-特里(BT)模型引入实例依赖的翻转概率。通过利用与偏好标注相关的特征,捕捉判断不确定性并建模翻转模式。实际中,设计了一种简单高效的迭代优化算法,兼容原始RLHF与DPO流程。实验在多种场景下评估了所提方法及基线方法,验证了其有效性。

原文摘要 · Abstract (English)

Quality of datasets plays an important role in large language model (LLM) alignment. In collecting human feedback, however, preference flipping is ubiquitous and causes corruption in data annotation; the issue necessitates the alignment algorithms with improved robustness against potential flipped pairs. To this end, this paper introduces a Flipping-Aware Direct Preference Optimization (FA-DPO) algorithm tailored to preference flipping from a reinforcement learning with human feedback (RLHF) perspective. We dissect the inherent human intention model and the preference flipping mechanism introduced by external factors as two distinct stages; in the latter, we introduce an instance-dependent flipping probability on the basis of the Bradley-Terry (BT) model. Further, by leveraging features relevant to preference annotation, we capture uncertainty in judgments and model preference flipping patterns. In practice, we design a simple yet efficient iterative optimization algorithm compatible with the original RLHF and DPO algorithms. In our experiments, we investigate the instance-dependent preference flipping model under multiple circumstances for evaluation of our proposed method, as well as other baseline methods.

强化学习偏好对齐鲁棒性人类反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。