通过不确定性估计提升机器人在人群中的安全导航泛化能力
Towards Generalizable Safety in Crowd Navigation via Conformal Uncertainty Handling
- 用自适应置信推断生成行人预测不确定性,融入强化学习
- 分布内成功率96.93%,碰撞减少3.72倍,轨迹侵入减少2.43倍
- 适用于真实机器人,对速度、行为、群体动态变化均具鲁棒性
移动机器人在人群环境中使用强化学习训练时,面对分布外场景常出现性能下降。我们提出,通过合理建模行人的不确定性,可让机器人学习到对分布偏移具有鲁棒性的安全导航策略。方法在智能体观测中加入由自适应置信推断生成的预测不确定性估计,并通过约束强化学习引导智能体行为。该系统能调节动作,增强对分布偏移的适应能力。在分布内场景下,成功率达96.93%,优于此前最优基线超8.80%,碰撞次数减少3.72倍,对真实人类未来轨迹的侵入减少2.43倍。在三种分布外场景(速度变化、策略改变、个体转群体)中,本方法展现出更强鲁棒性。已在真实机器人上部署,实验表明其在稀疏与密集人群互动中均能做出安全稳健决策。代码与视频见https://gen-safe-nav.github.io/。
原文摘要 · Abstract (English)
Mobile robots navigating in crowds trained using reinforcement learning are known to suffer performance degradation when faced with out-of-distribution scenarios. We propose that by properly accounting for the uncertainties of pedestrians, a robot can learn safe navigation policies that are robust to distribution shifts. Our method augments agent observations with prediction uncertainty estimates generated by adaptive conformal inference, and it uses these estimates to guide the agent's behavior through constrained reinforcement learning. The system helps regulate the agent's actions and enables it to adapt to distribution shifts. In the in-distribution setting, our approach achieves a 96.93% success rate, which is over 8.80% higher than the previous state-of-the-art baselines with over 3.72 times fewer collisions and 2.43 times fewer intrusions into ground-truth human future trajectories. In three out-of-distribution scenarios, our method shows much stronger robustness when facing distribution shifts in velocity variations, policy changes, and transitions from individual to group dynamics. We deploy our method on a real robot, and experiments show that the robot makes safe and robust decisions when interacting with both sparse and dense crowds. Our code and videos are available on https://gen-safe-nav.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。