用风险敏感策略和可达性验证,让机器人在复杂环境更安全导航。
Safety-Constrained Reinforcement Learning with Post-Training Reachability Verification for Robot Navigation

- 训练时用CVaR约束优化,关注高成本尾部风险而非平均表现。
- 测试后通过神经网络可达集分析,98.3%场景成功且安全率最高。
- 适合对安全性要求高的机器人导航任务,尤其看重形式化验证。
移动机器人在杂乱环境中导航需应对高后果感知不确定性下的可靠性挑战。现有多数安全强化学习方法仅以累计成本均值评估安全性,可能掩盖危险的尾部风险行为。为此,我们提出一种框架:基于离策略TD3骨干网络,采用条件风险价值(CVaR)约束优化训练风险敏感策略,并在训练后通过神经网络可达性验证评估安全裕度。训练中,策略在累积成本的CVaR约束下优化,增强对高成本尾部结果的敏感性。训练后,利用泰勒模型分析,在有界观测不确定性下计算动作可达集,得到安全率指标,量化策略在预设安全裕度内可到达状态的比例。关键发现是,CVaR训练策略在所有评估状态下保持更大的障碍物安全裕度,显著更易通过形式化可达性验证。十种导航场景、六种基线对比实验表明,本方法实现98.3%成功率,为所有方法中最高安全验证率;同时揭示平均成本排名与可达性安全排名存在分歧,说明可达性验证能捕捉经验成本指标遗漏的风险。进一步在物理Clearpath Jackal机器人上验证,成功实现从仿真到现实的迁移。
原文摘要 · Abstract (English)
Safe navigation for mobile robots demands policies that remain reliable under the high-consequence perception uncertainty of cluttered environments. Yet most existing safe reinforcement learning (RL) methods assess safety through average cumulative cost. Such metrics can mask dangerous tail-risk behaviors. To address this, we propose a framework that trains risk-sensitive policies through Conditional Value-at-Risk (CVaR) constrained optimization on an off-policy TD3 backbone and evaluates their safety margins post-training through neural network reachability verification. During training, the policy is optimized under CVaR constraints on cumulative costs, promoting sensitivity to high-cost tail outcomes rather than average behavior alone. After training, we compute action reachable sets under bounded observation uncertainty using Taylor Model analysis, yielding a safety rate metric that quantifies the proportion of evaluated states at which the policy's reachable action set remains within prescribed safety margins. A key finding is that policies trained with CVaR constraints maintain larger safety margins from obstacles across evaluated states. This makes them significantly more amenable to formal reachability verification. Experiments across ten navigation scenarios and six baselines show that our method achieves a 98.3\% success rate, the highest safety verification rate among all compared methods, while revealing that average cost rankings and reachability-based safety rankings can diverge. This indicates that reachability verification captures risks which are missed by empirical cost metrics alone. We further validate our approach on a physical Clearpath Jackal robot, demonstrating successful sim-to-real transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。