用分组对比可视化提升人类反馈强化学习效率
Interactive Groupwise Comparison for Reinforcement Learning from Human Feedback
- 通过分层聚类展示多样本行为,支持交互式分组比较
- 在6个机器人任务中使最终奖励提升69.34%,错误率更低
- 集成主动学习建议对比组,适合人机对齐研究者使用
基于人类反馈的强化学习(RLHF)已成为对齐人工智能行为与人类偏好的关键技术。传统方法依赖成对比较:由人类评分者判断两个样本中更优者。本文提出一种交互式可视化界面,充分利用人类视觉能力对多组样本进行整体比较与探索。该界面包含两个联动视图:1)探索视图,以分层聚类结构呈现所有采样行为的上下文概览;2)对比视图,展示用户选定的两组行为供评判。用户可在两视图间迭代操作,高效探索大规模行为集合。此外,我们设计了主动学习策略,推荐最优对比组。在六个模拟机器人任务中的评估表明,该方法使最终奖励提升69.34%,政策性能更优且错误率更低。代码已开源,可无缝集成至RLHF训练流程,支持人机对齐研究。
原文摘要 · Abstract (English)
Reinforcement learning from human feedback (RLHF) has emerged as a key enabling technology for aligning AI behaviour with human preferences. The traditional way to collect data in RLHF is via pairwise comparisons: human raters are asked to indicate which one of two samples they prefer. We present an interactive visualisation that better exploits the human visual ability to compare and explore whole groups of samples. The interface is comprised of two linked views: 1) an exploration view showing a contextual overview of all sampled behaviours organised in a hierarchical clustering structure; and 2) a comparison view displaying two selected groups of behaviours for user queries. Users can efficiently explore large sets of behaviours by iterating between these two views. Additionally, we devised an active learning approach suggesting groups for comparison. As shown by our evaluation in six simulated robotics tasks, our approach increases the final rewards by 69.34%. It leads to lower error rates and better policies. We open-source the code that can be easily integrated into the RLHF training loop, supporting research on human-AI alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。