用可证明安全的分割技术,让机器人在感知不全时也能安全导航。
Learning to Navigate Under Imperfect Perception: Conformalised Segmentation for Safe Reinforcement Learning
- 基于置信区间理论,对语义分割结果提供有限样本安全保证。
- 在卫星数据集上实现最高6倍的危险区域覆盖,误检率降低近50%。
- 适用于自动驾驶、无人机等对安全性要求极高的场景。
在安全关键环境中可靠导航需要精准的危险感知与规范化的不确定性处理,以强化下游安全决策。尽管现有方法有效,但通常假设危险检测能力完美,而现有的不确定性感知感知方法缺乏有限样本保障。本文提出COPPOL,一种基于置信区间的感知-策略学习框架,将分布无关、有限样本的安全性保证融入语义分割,生成校准后的危险区域地图,并提供漏检率的严格边界。这些地图用于构建风险感知的成本场,指导下游强化学习规划。在两个卫星衍生基准上,相比基线方法,COPPOL将危险区域覆盖提升至最多6倍,在几乎完全识别危险区域的同时,导航过程中的违规行为减少约50%。更重要的是,该方法在分布外场景下仍保持安全与效率鲁棒性。
原文摘要 · Abstract (English)
Reliable navigation in safety-critical environments requires both accurate hazard perception and principled uncertainty handling to strengthen downstream safety handling. Despite the effectiveness of existing approaches, they assume perfect hazard detection capabilities, while uncertainty-aware perception approaches lack finite-sample guarantees. We present COPPOL, a conformal-driven perception-to-policy learning approach that integrates distribution-free, finite-sample safety guarantees into semantic segmentation, yielding calibrated hazard maps with rigorous bounds for missed detections. These maps induce risk-aware cost fields for downstream RL planning. Across two satellite-derived benchmarks, COPPOL increases hazard coverage (up to 6x) compared to comparative baselines, achieving near-complete detection of unsafe regions while reducing hazardous violations during navigation (up to approx 50%). More importantly, our approach remains robust to distributional shift, preserving both safety and efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。