通过干预人类表达偏好方式,提升其与强化学习对齐模型的契合度。
Influencing Humans to Conform to Preference Models for RLHF
- 通过展示奖励函数隐含量、训练匹配模型、改写提问方式来影响人类偏好表达。
- 三种干预均显著提升人类偏好与目标模型的一致性,改善对齐效果。
- 适合关注人类反馈对齐、提示工程与人机协同设计的研究者。
设计基于人类反馈的强化学习(RLHF)算法以逼近人类不可观测的奖励函数,需隐含或显式假设一个关于人类偏好的模型。若偏好模型无法准确描述人类生成偏好的机制,将导致对人类奖励函数的近似效果不佳。本文开展三项人类实验,评估是否可通过干预使真实人类偏好更贴合期望的偏好模型。关键在于:不改变人类内在奖励函数,而是调整其使用该函数生成偏好的方式,使其更符合特定RLHF算法所依赖的偏好模型假设。我们提出三种干预策略:向人类展示偏好模型背后的隐含数量(通常不可观测);训练人类遵循特定偏好模型;修改偏好获取的问题形式。所有干预均产生显著效果,为提升偏好数据质量及后续学习奖励函数的对齐性提供了实用工具。本文确立了一个新的研究方向:设计界面与训练干预,以增强人类对算法建模假设的符合度。
原文摘要 · Abstract (English)
Designing a reinforcement learning from human feedback (RLHF) algorithm to approximate a human's unobservable reward function requires assuming, implicitly or explicitly, a model of human preferences. A preference model that poorly describes how humans generate preferences risks learning a poor approximation of the human's reward function. In this paper, we conduct three human studies to asses whether one can influence the expression of real human preferences to more closely conform to a desired preference model. Importantly, our approach does not seek to alter the human's unobserved reward function. Rather, we change how humans use this reward function to generate preferences, such that they better match whatever preference model is assumed by a particular RLHF algorithm. We introduce three interventions: showing humans the quantities that underlie a preference model, which is normally unobservable information derived from the reward function; training people to follow a specific preference model; and modifying the preference elicitation question. All intervention types show significant effects, providing practical tools to improve preference data quality and the resultant alignment of the learned reward functions. Overall we establish a novel research direction in model alignment: designing interfaces and training interventions to increase human conformance with the modeling assumptions of the algorithm that will learn from their input.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。