用对称损失提升偏好噪声下的策略优化鲁棒性
On Symmetric Losses for Robust Policy Optimization with Noisy Preferences
- 将奖励建模视为分类问题,引入对称损失增强抗噪能力
- 在噪声标签下仍能保持排序一致性,确保策略有效提升
- 适合存在标注偏差的真实场景,如人类偏好反馈
基于人类偏好的策略优化是使语言模型对齐人类意图的关键。本文聚焦强化学习中的人类反馈(RLHF)和离线偏好优化中的核心环节——奖励建模。传统方法通常假设标注准确,但现实偏好数据常因人为错误或偏见而含噪声。本文提出一种在噪声偏好下稳健优化策略的理论框架,将奖励建模视为分类任务,从而可应用对称损失(symmetric losses),因其在分类中对标签噪声具有鲁棒性。由此提出的对称偏好优化(SymPO)方法,在理论上证明即使在噪声标签下,所生成的奖励仍保持排序一致性(rank-preserving),而该性质足以支持策略改进。在合成与真实任务上的实验验证了SymPO的有效性。
原文摘要 · Abstract (English)
Optimizing policies based on human preferences is key to aligning language models with human intent. This work focuses on reward modeling, a core component in reinforcement learning from human feedback (RLHF), and offline preference optimization, such as direct preference optimization. Conventional approaches typically assume accurate annotations. However, real-world preference data often contains noise due to human errors or biases. We propose a principled framework for robust policy optimization under noisy preferences, viewing reward modeling as a classification problem. This allows us to leverage symmetric losses, known for their robustness to label noise in classification, leading to our Symmetric Preference Optimization (SymPO) method. We prove that symmetric losses enable successful policy optimization even under noisy labels, as the resulting reward remains rank-preserving -- a property sufficient for policy improvement. Experiments on synthetic and real-world tasks demonstrate the effectiveness of SymPO.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。