用人类示范数据自动提取隐式奖励,提升AI对齐效果
Inverse RL Helps Align AI by Imitating Humans
- 从人类示范中逆向学习隐式奖励函数,无需额外标注
- 在不依赖监督损失情况下显著提升基础模型表现
- 可适配不同用户偏好,实现个性化AI对齐
语言模型对齐旨在使模型行为稳定反映有益性、安全性及指令遵循等理想特性。现有方法多依赖示范数据的监督微调或基于验证器/人类反馈的强化学习奖励。但一个重要问题仍未解决:仅凭示范能否生成可观察、可复用、可在线优化的隐式奖励?受逆强化学习启发,我们提出基于示范估计的投影对齐奖励(PARED)。PARED通过轻量级判别器,在响应级特征空间中区分示范样本与策略自采样,显式恢复出隐含奖励函数。与标准奖励模型不同,PARED无需任务特定偏好标注:示范提供任务监督,可叠加AI反馈作为额外监督维度。实验表明,该恢复奖励在推理时重排序与对抗性在线强化学习中均能改进基础策略,且在标准监督微调后进一步提升性能。此外,我们证明PARED可用于情境化对齐,使单一策略适应不同受众偏好。
原文摘要 · Abstract (English)
Language model alignment aims to make model behavior reliably reflect desirable properties such as helpfulness, safety, and instruction following. Current approaches typically use supervised fine-tuning on demonstrations or reinforcement learning with rewards derived from verifiers or human feedback. These paradigms leave an important question underexplored: can demonstrations alone yield an implicit reward that can be inspected, reused, and optimized on-policy to align AI? Motivated by inverse reinforcement learning, we introduce Projected Alignment Reward Estimated from Demonstrations (PARED). PARED recovers the implicit reward underlying expert demonstrations as an explicit function over a small set of response-level features, learned by a lightweight discriminator that separates demonstrations from the policy's own samples in this feature space. Unlike a standard reward model, PARED requires no task-specific preference annotations: demonstrations provide the task-specific supervision, which can be augmented with AI feedback as additional dimensions of supervision. Through experiments involving inference-time reranking and adversarial on-policy RL, we show that the recovered reward improves a base policy without a supervised loss and yields further gains when optimized after standard supervised fine-tuning. Additionally, we demonstrate that PARED can be used for contextual alignment, in which a single policy can be tailored to the preferences of different audiences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。