用人类注视信息提升强化学习对齐效率,加快收敛并降低成本。
Enhancing RLHF with Human Gaze Modeling
- 引入注视信息构建奖励模型,优化策略学习
- 在词元级稀疏奖励中使用注视数据,加速收敛
- 适合关注高效对齐与人机交互的研学者
基于人类反馈的强化学习(RLHF)能将语言模型对齐人类偏好,但计算成本高。本文探索两种利用人类注视建模增强RLHF的方法:(1) 注视感知的奖励模型;(2) 在词元级别基于注视分布稀疏奖励。实验表明,引入注视信息的RLHF实现更快收敛,同时保持或轻微提升性能,从而降低策略优化阶段的计算开销。结果表明,人类注视是政策优化中一个有价值且未被充分利用的信号,为提升RLHF效率指明了有前景的方向。
原文摘要 · Abstract (English)
Reinforcement Learning from Human Feedback (RLHF) aligns language models with human preferences but is computationally expensive. We explore two approaches that leverage human gaze modeling to enhance RLHF: (1) gaze-aware reward models and (2) gaze-based distribution of sparse rewards at token level. Our experiments demonstate that gaze-informed RLHF achieves faster convergence while maintaining or slightly improving performance, thus, reducing computational costs during policy optimization. These results show that human gaze provides a valuable and underused signal for policy optimization, pointing to a promising direction for improving RLHF efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。