解决大模型对齐中内容相关噪声干扰问题,提升真实偏好学习鲁棒性。
One Goal, Many Challenges: Robust Preference Optimization Amid Content-Aware and Multi-Source Noise
- 通过多目标优化分离真实偏好与内容相关噪声
- 在合成数据集上显著降低响应长度、有害性等次要偏见影响
- 利用后门攻击机制高效建模并控制多种噪声源,适合高可靠性对齐场景
大语言模型在生成类人回复方面取得显著进展,主要得益于偏好对齐技术。然而,现有方法通常假设人类反馈无偏,这在真实场景中极少成立。本文提出内容感知噪声鲁棒偏好优化(CNRPO),一种新型框架,用于应对偏好学习中多种内容依赖性噪声。CNRPO采用多目标优化策略,将真实偏好与内容相关噪声分离,有效缓解其影响。我们利用后门攻击机制,在单一模型中高效学习和控制各类噪声源。理论分析与在多个合成噪声数据集上的广泛实验表明,CNRPO能显著提升对主干人类偏好的对齐效果,同时有效控制响应长度、有害性等次级噪声与偏见。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have made significant strides in generating human-like responses, largely due to preference alignment techniques. However, these methods often assume unbiased human feedback, which is rarely the case in real-world scenarios. This paper introduces Content-Aware Noise-Resilient Preference Optimization (CNRPO), a novel framework that addresses multiple sources of content-dependent noise in preference learning. CNRPO employs a multi-objective optimization approach to separate true preferences from content-aware noises, effectively mitigating their impact. We leverage backdoor attack mechanisms to efficiently learn and control various noise sources within a single model. Theoretical analysis and extensive experiments on different synthetic noisy datasets demonstrate that CNRPO significantly improves alignment with primary human preferences while controlling for secondary noises and biases, such as response length and harmfulness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。