研究噪声反馈对模型对齐效果的影响,揭示真实场景下训练的局限性。
How Well Can Preference Optimization Generalize Under Noisy Feedback?
- 分析不同噪声类型下偏好优化的泛化性能衰减规律
- 发现噪声率越高,模型对齐效果越差,且与数据分布相关
- 适用于DPO、IPO等主流对齐方法,适合关注模型可靠性的人看
随着大语言模型能力提升,使其与人类偏好对齐变得至关重要。偏好优化通过人类反馈区分优选与非优选响应,成为对齐的核心手段。然而,现有研究多假设反馈无噪声,这与人类判断固有的错误和不一致不符。本文研究噪声反馈对偏好优化的影响,提供在噪声条件下的泛化保证。考虑了误标和不确定性等现实噪声来源,不同于传统收敛假设,聚焦有限步训练,更贴近实际。分析表明,泛化性能随噪声类型和噪声率变化,受偏好数据分布及样本数量影响。该理论适用于DPO、IPO、SLiC等多种偏好优化损失函数。在当代大模型上的实证验证了结论的实用性,为构建真正符合人类偏好的AI系统提供了关键洞见。
原文摘要 · Abstract (English)
As large language models (LLMs) advance their capabilities, aligning these models with human preferences has become crucial. Preference optimization, which trains models to distinguish between preferred and non-preferred responses based on human feedback, has become a crucial component for aligning LLMs. However, most existing works assume noise-free feedback, which is unrealistic due to the inherent errors and inconsistencies in human judgments. This paper addresses the impact of noisy feedback on preference optimization, providing generalization guarantees under these conditions. In particular, we consider noise models that correspond to common real-world sources of noise, such as mislabeling and uncertainty. Unlike traditional analyses that assume convergence, our work focuses on finite-step preference optimization, offering new insights that are more aligned with practical LLM training. We describe how generalization decays with different types of noise across levels of noise rates based on the preference data distribution and number of samples. Our analysis for noisy preference learning applies to a broad family of preference optimization losses such as DPO, IPO, SLiC, etc. Empirical validation on contemporary LLMs confirms the practical relevance of our findings, offering valuable insights for developing AI systems that align with human preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。