用群体方法提升人类偏好下的强化学习探索能力
PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning
- 采用多智能体群体保持多样性,更好探索偏好空间
- 在复杂环境和人类错误标注下仍表现稳定,减少反馈依赖
- 适合需要真实人类反馈的智能系统训练场景
基于人类偏好的强化学习(PbRL)无需预设奖励函数即可学习行为,但现有方法常因偏好空间探索不足而过早收敛至次优策略,仅满足部分偏好。本文通过群体方法解决该问题:保持多个智能体构成的种群,实现更全面的偏好空间探索。多样性使智能体产生差异显著的行为,便于人类清晰区分并提供有效反馈,这对现实场景至关重要。实验表明,当前方法易陷入局部最优、需大量反馈,且在人类误判相似轨迹时性能骤降;而本方法在教师错误标注下依然稳健,尤其在复杂奖励环境下偏好探索能力显著增强。
原文摘要 · Abstract (English)
Preference-based reinforcement learning (PbRL) has emerged as a promising approach for learning behaviors from human feedback without predefined reward functions. However, current PbRL methods face a critical challenge in effectively exploring the preference space, often converging prematurely to suboptimal policies that satisfy only a narrow subset of human preferences. In this work, we identify and address this preference exploration problem through population-based methods. We demonstrate that maintaining a diverse population of agents enables more comprehensive exploration of the preference landscape compared to single-agent approaches. Crucially, this diversity improves reward model learning by generating preference queries with clearly distinguishable behaviors, a key factor in real-world scenarios where humans must easily differentiate between options to provide meaningful feedback. Our experiments reveal that current methods may fail by getting stuck in local optima, requiring excessive feedback, or degrading significantly when human evaluators make errors on similar trajectories, a realistic scenario often overlooked by methods relying on perfect oracle teachers. Our population-based approach demonstrates robust performance when teachers mislabel similar trajectory segments and shows significantly enhanced preference exploration capabilities,particularly in environments with complex reward landscapes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。