用少量偏好示例实时适配奖励模型,让AI更好理解多样人类偏好。
In-Context Reward Adaptation for Robust Preference Modeling

- 基于Transformer在上下文中动态推断奖励结构,无需重训练。
- 引入人类反应时间作为辅助信号,显著提升对未知偏好的适应能力。
- 适合需要灵活应对不同人群偏好的人机对齐场景。
强化学习从人类反馈(RLHF)通常依赖静态奖励模型来对齐大语言模型与人类偏好。然而人类价值观本质上多样且异质,单一奖励模型往往难以泛化到未见的偏好领域。现有多种奖励框架虽试图解决此问题,但常受限于已知领域集合,无法在不进行昂贵重训的情况下适应未见的人类分布。本文提出「上下文奖励自适应」(In-Context Reward Adaptation),一种基于Transformer的框架,可即时建模多样且未见的人类偏好。通过利用Transformer的上下文学习能力,该方法从少量偏好示范中自适应推断底层奖励结构。我们发现标准Transformer存在对真实奖励的渐近偏差;而引入人类反应时间作为辅助输入信号后,模型成功适应了此前未见领域的偏好。结果表明,该方法为偏好建模提供了更鲁棒的基础,能有效表示异质奖励与偏好分布变化,并为更灵活的人机对齐提供可扩展路径。
原文摘要 · Abstract (English)
Reinforcement Learning from Human Feedback (RLHF) typically relies on static reward models to align Large Language Models with human preferences. However, human values are inherently diverse and heterogeneous, and a single reward model often lacks the robustness required to generalize to unseen preference domains. While existing multi-reward frameworks attempt to address this, they are often restricted to a fixed set of known domains and fail to adapt to unseen human distributions without costly retraining. In this work, we propose In-Context Reward Adaptation, a transformer-based framework designed to model diverse and unseen human preferences on the fly. By leveraging the in-context learning capabilities of transformers, our approach adaptively infers the underlying reward structure from a small set of preference demonstrations. We demonstrate that while a standard transformer architecture is insufficient for this task by characterizing an asymptotic bias to the ground-truth, incorporating human response time as an auxiliary input signal enables the model to successfully adapt to preferences from previously unseen domains. Our findings show that this approach provides a more robust foundation for preference modeling, allowing for the representation of heterogeneous rewards and preference distribution shift, and offering a scalable path toward more flexible human-AI alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。