用偏好学习提升实时人类反馈下的持续强化学习性能
Pref-GUIDE: Continual Policy Learning from Real-Time Human Feedback via Preference-Based Learning
- 将实时标量反馈转化为偏好数据,提升奖励模型学习质量
- 通过窗口内行为对比和用户投票,减少反馈噪声与不一致
- 在三个挑战性环境中超越标量反馈基线,接近专家设计奖励
在任务目标难以通过密集奖励函数定义时,基于人类反馈训练强化学习智能体至关重要。以往方法依赖离线轨迹比较获取人类偏好,但在在线学习场景中无法使用。近期方法尝试收集实时标量反馈以指导智能体行为并训练奖励模型,但标量反馈常存在噪声与不一致性,限制了学习奖励的准确性与泛化能力。本文提出 Pref-GUIDE 框架,将实时标量反馈转化为偏好数据,以改进持续策略训练中的奖励模型学习。Pref-GUIDE Individual 通过短时间窗口内行为对比并过滤模糊反馈,缓解时间不一致性;Pref-GUIDE Voting 进一步通过聚合多用户群体反馈形成共识偏好,增强鲁棒性。在三个挑战性环境中,Pref-GUIDE 显著优于标量反馈基线,其投票变体甚至超越专家设计的密集奖励。通过将标量反馈重构为带群体共识的结构化偏好,Pref-GUIDE 为在线强化学习中高效利用人类输入提供了一种可扩展且原则性强的方法。
原文摘要 · Abstract (English)
Training reinforcement learning agents with human feedback is crucial when task objectives are difficult to specify through dense reward functions. While prior methods rely on offline trajectory comparisons to elicit human preferences, such data is unavailable in online learning scenarios where agents must adapt on the fly. Recent approaches address this by collecting real-time scalar feedback to guide agent behavior and train reward models for continued learning after human feedback becomes unavailable. However, scalar feedback is often noisy and inconsistent, limiting the accuracy and generalization of learned rewards. We propose Pref-GUIDE, a framework that transforms real-time scalar feedback into preference-based data to improve reward model learning for continual policy training. Pref-GUIDE Individual mitigates temporal inconsistency by comparing agent behaviors within short windows and filtering ambiguous feedback. Pref-GUIDE Voting further enhances robustness by aggregating reward models across a population of users to form consensus preferences. Across three challenging environments, Pref-GUIDE significantly outperforms scalar-feedback baselines, with the voting variant exceeding even expert-designed dense rewards. By reframing scalar feedback as structured preferences with population feedback, Pref-GUIDE offers a scalable and principled approach for harnessing human input in online reinforcement learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。