arXiv:2608.27793cs.RO2026-08

用视觉语言模型导航水下洞穴,靠光线和结构判断安全路径。

CAVE-NAV: VLM-Based Autonomous 3D Navigation in Underwater Cave Environments

论文配图:CAVE-NAV: VLM-Based Autonomous 3D Navigation in Underwater Cave Environments
图 1 · 摘自论文原文
  • 结合图像、深度图和声呐数据,用VLM推理可通行方向
  • 模拟测试中全程无碰撞,保持与洞壁安全距离
  • 适合水下搜救与科研探索,无需实时人工干预

水下洞穴的自主导航对搜救、科学探测和紧急撤离至关重要。传统系统依赖密集视觉特征进行定位与建图,但在水下洞穴中,视觉退化会削弱特征定位效果,声呐建图可能过于保守,通信限制也使实时人工指导不可行。为此,我们提出一种基于视觉语言模型(VLM)与思维链(CoT)推理的自主水下洞穴导航框架,通过多模态输入(包括RGB图像、深度图和声呐测得的垂直净空)感知环境线索,如光强梯度、通道形态和几何复杂度,推断可通行方向,实现安全的三维导航。在多种洞穴拓扑的高保真仿真中,该框架成功完成所有端到端穿越任务,无碰撞且保持与洞壁的安全距离。

原文摘要 · Abstract (English)

Autonomous navigation in underwater cave environments is essential for search-and-rescue operations, scientific exploration, and emergency egress. Traditional navigation systems commonly depend on dense visual features for localization and mapping. In underwater caves, however, visual degradation can undermine feature-based localization, sonar-based mapping may yield overly conservative obstacle representations, and communication constraints preclude real-time human guidance. To address these limitations, we propose an autonomous underwater cave navigation framework that leverages a vision-language model (VLM) with Chain-of-Thought (CoT) reasoning to infer navigable directions from environmental cues, including light intensity gradients, passage morphology, and geometric complexity, captured through multimodal inputs comprising RGB imagery, depth maps, and sonar-based vertical-clearance measurements, thereby supporting safe 3D navigation through confined cave passages. High-fidelity simulations across multiple cave topologies demonstrate that the proposed framework completes all evaluated end-to-end traversals without collisions while maintaining safe clearance from cave boundaries.

水下导航视觉语言模型多模态融合3D路径规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。