将人类偏好从数值奖励扩展为结构化表示,提升可解释性与应用广度。
LRHP: Learning Representations for Human Preferences via Preference Pairs
- 通过偏好对学习人类偏好的结构化表示,超越传统奖励建模
- 在偏好数据筛选与边际预测任务中显著优于基线方法
- 适合需要理解偏好机制的AI对齐与个性化系统研究者
为提升人类偏好对齐训练效果,现有研究构建了大量由偏好对组成的标注数据集,通常用于通过奖励建模将人类偏好编码为单一数值,作为强化学习中的奖励信号。然而,将偏好简化为数值会限制其分析深度,并限制其在非RLHF场景中的应用。本文提出一种偏好表示学习任务,旨在构建更丰富、更结构化的偏好表示。我们进一步开发了一个通用框架——基于偏好对的人类偏好表示学习(LRHP),突破传统奖励建模局限。我们在两个下游任务中验证了该表示的有效性:偏好数据选择和偏好边际预测。基于学习到的偏好表示,模型在两项任务中均表现优异,显著超越基线方法。
原文摘要 · Abstract (English)
To improve human-preference alignment training, current research has developed numerous preference datasets consisting of preference pairs labeled as "preferred" or "dispreferred". These preference pairs are typically used to encode human preferences into a single numerical value through reward modeling, which acts as a reward signal during reinforcement learning from human feedback (RLHF). However, representing these human preferences as a numerical value complicates the analysis of these preferences and restricts their broader applications other than RLHF. In contrast, in this work, we introduce a preference representation learning task that aims to construct a richer and more structured representation of human preferences. We further develop a more generalizable framework, Learning Representations for Human Preferences via preference pairs (namely LRHP), which extends beyond traditional reward modeling to tackle this task. We verify the utility of preference representations in two downstream tasks: preference data selection and preference margin prediction. Building upon the human preferences in representations, we achieve strong performance in both tasks, significantly outperforming baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。