用主动查询选关键动作,让众包反馈更高效
Active Query Selection for Crowd-Based Reinforcement Learning
- 结合众包标注噪声建模与主动学习,智能选需反馈的动作
- 在血糖控制任务中,比基线更快收敛且效果更优
- 适合专家少、反馈成本高的真实场景应用
基于偏好的强化学习在奖励信号难定义或与人类意图不一致的环境中日益重要。但其效果常受限于可靠人工反馈成本高、获取难,尤其在专家稀缺或错误代价高的领域。为此,我们提出新框架:融合概率众包建模以处理多标注者噪声,结合主动学习优先请求最具信息量的动作反馈。扩展Advise算法支持多训练者,实时估计标注者可靠性,并采用基于熵的查询选择策略引导反馈请求。在涵盖合成与现实启发式环境的任务中评估,包括2D游戏(Taxi、Pacman、Frozen Lake)及使用临床认可的UVA/Padova模拟器的1型糖尿病血糖控制任务。初步结果表明,对不确定轨迹进行反馈训练的智能体在多数任务中学习速度更快,且在血糖控制任务上优于基线。
原文摘要 · Abstract (English)
Preference-based reinforcement learning has gained prominence as a strategy for training agents in environments where the reward signal is difficult to specify or misaligned with human intent. However, its effectiveness is often limited by the high cost and low availability of reliable human input, especially in domains where expert feedback is scarce or errors are costly. To address this, we propose a novel framework that combines two complementary strategies: probabilistic crowd modelling to handle noisy, multi-annotator feedback, and active learning to prioritize feedback on the most informative agent actions. We extend the Advise algorithm to support multiple trainers, estimate their reliability online, and incorporate entropy-based query selection to guide feedback requests. We evaluate our approach in a set of environments that span both synthetic and real-world-inspired settings, including 2D games (Taxi, Pacman, Frozen Lake) and a blood glucose control task for Type 1 Diabetes using the clinically approved UVA/Padova simulator. Our preliminary results demonstrate that agents trained with feedback on uncertain trajectories exhibit faster learning in most tasks, and we outperform the baselines for the blood glucose control task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。