通过人类干预预测未来行为,提升智能体安全学习效率
Predictive Preference Learning from Human Interventions

- 将人类干预扩展至未来L步,构建偏好信号预测机制
- 在自动驾驶与机器人操控中减少70%以上的人工示范需求
- 适合需要高安全性、低人工标注的强化学习场景
学习人类参与旨在引入人类主体监控并纠正智能体的行为错误。尽管大多数交互式模仿学习方法聚焦于当前状态的动作修正,却未调整未来状态下的动作,可能带来潜在风险。为此,我们提出从人类干预中进行预测性偏好学习(PPL),利用人类干预中隐含的偏好信号来预测未来轨迹。PPL的核心思想是将每次人类干预向后扩展L个未来时间步,称为偏好视野,假设在该视野内代理执行相同动作、人类作出相同干预。通过对这些未来状态进行偏好优化,专家修正被传播至智能体预期探索的安全敏感区域,显著提升学习效率并减少所需的人类示范。我们在自动驾驶与机器人操作基准上进行了实验,验证了该方法的高效性与通用性。理论分析进一步表明,选择合适的偏好视野长度L可在覆盖危险状态与标签正确性之间取得平衡,从而限制算法最优性差距。演示与代码已公开:https://metadriverse.github.io/ppl
原文摘要 · Abstract (English)
Learning from human involvement aims to incorporate the human subject to monitor and correct agent behavior errors. Although most interactive imitation learning methods focus on correcting the agent's action at the current state, they do not adjust its actions in future states, which may be potentially more hazardous. To address this, we introduce Predictive Preference Learning from Human Interventions (PPL), which leverages the implicit preference signals contained in human interventions to inform predictions of future rollouts. The key idea of PPL is to bootstrap each human intervention into L future time steps, called the preference horizon, with the assumption that the agent follows the same action and the human makes the same intervention in the preference horizon. By applying preference optimization on these future states, expert corrections are propagated into the safety-critical regions where the agent is expected to explore, significantly improving learning efficiency and reducing human demonstrations needed. We evaluate our approach with experiments on both autonomous driving and robotic manipulation benchmarks and demonstrate its efficiency and generality. Our theoretical analysis further shows that selecting an appropriate preference horizon L balances coverage of risky states with label correctness, thereby bounding the algorithmic optimality gap. Demo and code are available at: https://metadriverse.github.io/ppl
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。