让无人潜航器在看不清环境时也能安全导航,靠的是融合感知与强化学习的统一框架。
Unified Planning-Learning Framework for Robust UUV Navigation Under Partial Observability

- 仅用声呐和深度图构建地图,结合全局规划与强化学习实现长程路径与短程避障
- 在动态环境下比传统方法更安全,降低碰撞风险,提升导航鲁棒性
- 适合需要高可靠性水下自主导航的科研或工程场景
本文提出一种仅依赖观测的无人潜航器(UUV)自主导航框架,适用于动态水下环境。该框架仅利用机载声呐和深度图像构建持续占用地图,采用带清空约束的全局规划器(GP)提供长程结构指引,并集成强化学习(RL)策略处理短距离追踪与实时避障。为应对部分可观测性,系统从传感器数据中学习紧凑的隐状态表示,编码环境结构、障碍物动态及不确定性。引入分阶段监督的行为树(BT)蒸馏方法提升安全性与训练稳定性,同时通过在线隐模型不确定性校准蒸馏权重,强化学习中重点关注不确定区域。时间到碰撞(TTC)和清空提示作为显式特征融入规划与局部策略。通过基于NVIDIA Isaac Sim的高保真GPU加速仿真,建立可复现的多种子评估协议,与纯行为树和标准强化学习基线对比。结果表明,该框架在动态条件下显著提升了鲁棒性与安全性,提供了一个统一的混合规划-学习架构和可复现的方法论,支持部分可观测下的鲁棒UUV自主导航。
原文摘要 · Abstract (English)
This paper presents an observation-only autonomy framework for Unmanned Underwater Vehicles (UUVs) navigation in dynamic underwater environments that integrates persistent occupancy mapping, global clearance-aware planning, and risk-aware local control. The proposed pipeline constructs occupancy maps solely from onboard sonar and depth image observations, adapts a clearance-constrained global planner (GP) to provide long-horizon structure, and integrates a reinforcement learning (RL) policy to handle short-range tracking and reactive avoidance. To further support decision-making under partial observability, the system learns a compact latent state representation from onboard sensor data, encoding environmental structure, obstacle dynamics, and uncertainty. Behavior tree (BT) distillation with staged supervision is introduced to improve safety and training stability, while an uncertainty-calibrated distillation mechanism reweights teacher guidance using online latent-model uncertainty, emphasizing uncertain regimes during learning, with time-to-collision (TTC) and clearance cues remaining explicit in planning and local policy features. To demonstrate the efficacy of the framework, a reproducible multi-seed evaluation protocol is established in high-fidelity GPU-accelerated simulation using NVIDIA Isaac Sim, and performance is benchmarked against BT-only and standard RL baselines. The results obtained demonstrate improved robustness and safety under dynamic conditions, thus providing a general pipeline with a unified hybrid planning learning architecture and a reproducible methodology for robust UUV autonomy under partial observability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。