通过表征学习消除奖励模型中的虚假相关,提升对齐可靠性
Debiasing Reward Models by Representation Learning with Guarantees
- 从数据生成过程建模,分离真实偏好与虚假信号
- 无需虚假变量代理即可理论识别真实偏好因子
- 在合成与真实数据上验证了对虚假相关性的有效抑制
近年来,基于人类反馈的强化学习等对齐技术广泛用于将大语言模型与人类偏好对齐,其核心是训练和使用奖励模型。然而,这些模型常利用虚假相关性(如回应长度、歧视性表述、谄媚倾向、概念偏差)来预测偏好,问题日益突出。本文提出一个理论严谨的框架,旨在缓解奖励模型中的此类偏差,同时保留反映真实意图偏好的底层因素。我们首先构建数据生成过程的数学形式,假设观测数据(如文本)由虚假和非虚假潜在变量共同生成。研究发现,即便没有虚假变量的代理,非虚假潜在变量仍可从数据中理论识别。这一发现启发了一种实用方法:采用变分推断恢复这些变量,并用其训练更鲁棒的奖励模型。在合成与真实世界数据集上的实验表明,该方法能有效缓解虚假相关问题,显著提升奖励模型的稳定性与可靠性。
原文摘要 · Abstract (English)
Recent alignment techniques, such as reinforcement learning from human feedback, have been widely adopted to align large language models with human preferences by learning and leveraging reward models. In practice, these models often exploit spurious correlations, involving, e.g., response length, discrimination, sycophancy, and conceptual bias, which is a problem that has received increasing attention. In this work, we propose a principled framework that mitigates these biases in reward models while preserving the underlying factors that reflect intended preferences. We first provide a formulation of the data-generating process, assuming that the observed data (e.g., text) is generated from both spurious and non-spurious latent variables. We show that, interestingly, these non-spurious latent variables can be theoretically identified from data, regardless of whether a surrogate for the spurious latent variables is available. This further inspires a practical method that uses variational inference to recover these variables and leverages them to train reward models. Experiments on synthetic and real-world datasets demonstrate that our method effectively mitigates spurious correlation issues and yields more robust reward models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。