让AI通过反思对话学习用户个性化偏好,提升对多元价值观的对齐效果。
Reflective Verbal Reward Design for Pluralistic Alignment
- 用语言模型引导用户进行反思性对话,生成个性化偏好数据。
- 在30人实验中,准确率比非反思模型提升9-12%。
- 适合需要尊重个体差异价值判断的AI交互场景。
AI代理通常通过人类反馈强化学习(RLHF)与“人类价值观”对齐,即从聚合的人类反馈中训练单一奖励模型来指导行为。然而,人类价值观并非同质——不同人持有不同甚至冲突的价值观。将反馈聚合为单一模型可能过度压制少数群体偏好。为此,我们提出一种新型奖励建模方法,用于学习个体化奖励模型。该方法利用语言模型引导用户进行反思性对话,让用户批判代理行为并构建自身偏好。这一个性化对话历史(包含用户反思和被批注的示例)作为上下文,输入另一语言模型,形成个体化的奖励函数(我们称之为“言语奖励模型”),用于评估新轨迹。在30名参与者的实验中,该方法相比非反思言语奖励模型准确率提升9-12%,且比传统监督学习更高效。
原文摘要 · Abstract (English)
AI agents are commonly aligned with "human values" through reinforcement learning from human feedback (RLHF), where a single reward model is learned from aggregated human feedback and used to align an agent's behavior. However, human values are not homogeneous--different people hold distinct and sometimes conflicting values. Aggregating feedback into a single reward model risks disproportionately suppressing minority preferences. To address this, we present a novel reward modeling approach for learning individualized reward models. Our approach uses a language model to guide users through reflective dialogues where they critique agent behavior and construct their preferences. This personalized dialogue history, containing the user's reflections and critiqued examples, is then used as context for another language model that serves as an individualized reward function (what we call a "verbal reward model") for evaluating new trajectories. In studies with 30 participants, our method achieved a 9-12% improvement in accuracy over non-reflective verbal reward models while being more sample efficient than traditional supervised learning methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。