用深度强化学习让机器人在人流中安全导航,自动识别何时该保守避障。
Disentangling Uncertainty for Safe Social Navigation using Deep Reinforcement Learning
- 引入观测依赖方差与丢弃法,提升策略不确定性估计能力。
- 蒙特卡洛丢弃法对扰动更敏感,能更好区分不同类型的不确定性。
- 在复杂场景中减少碰撞,适合需安全交互的移动机器人应用。
自主移动机器人越来越多地应用于行人密集环境,安全导航与恰当的人类交互至关重要。尽管深度强化学习(DRL)可实现社交化机器人行为,但在新场景或受扰动环境下,难以判断策略是否不确定。决策中的未知不确定性可能导致碰撞或引发人类不适,是实现安全、风险感知导航的主要挑战。本文提出一种新方法,将偶然性、认知性和预测性不确定性估计集成至DRL导航框架中,用于策略分布的不确定性建模。通过在近端策略优化(PPO)算法中引入观测依赖方差(ODV)和丢弃法,并对比深度集成与蒙特卡洛丢弃(MC-dropout)在不同扰动下的不确定性估计能力。针对不确定决策情况,提出将机器人社交行为转为保守避障。实验表明,加入ODV与丢弃法后训练性能提升,且训练场景影响泛化效果;其中MC-dropout对扰动更敏感,能更准确关联不确定性类型与扰动源。结合安全动作选择,机器人可在受扰环境中显著降低碰撞率。
原文摘要 · Abstract (English)
Autonomous mobile robots are increasingly used in pedestrian-rich environments where safe navigation and appropriate human interaction are crucial. While Deep Reinforcement Learning (DRL) enables socially integrated robot behavior, challenges persist in novel or perturbed scenarios to indicate when and why the policy is uncertain. Unknown uncertainty in decision-making can lead to collisions or human discomfort and is one reason why safe and risk-aware navigation is still an open problem. This work introduces a novel approach that integrates aleatoric, epistemic, and predictive uncertainty estimation into a DRL navigation framework for policy distribution uncertainty estimates. We, therefore, incorporate Observation-Dependent Variance (ODV) and dropout into the Proximal Policy Optimization (PPO) algorithm. For different types of perturbations, we compare the ability of deep ensembles and Monte-Carlo dropout (MC-dropout) to estimate the uncertainties of the policy. In uncertain decision-making situations, we propose to change the robot's social behavior to conservative collision avoidance. The results show improved training performance with ODV and dropout in PPO and reveal that the training scenario has an impact on the generalization. In addition, MC-dropout is more sensitive to perturbations and correlates the uncertainty type to the perturbation better. With the safe action selection, the robot can navigate in perturbed environments with fewer collisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。