arXiv:2603.14786cs.RO2026-03被引 1

用分层框架让水下机器人智能导航,提升监测效率与安全性。

CORAL: COntextual Reasoning And Local Planning in A Hierarchical VLM Framework for Underwater Monitoring

  • 高阶语义推理与低阶控制解耦,由VLM选路线点,规划器避障执行
  • 覆盖率提升17.85%,碰撞减少100%,VLM调用次数降低57%
  • 适合需要长期自主水下巡检的生态监测场景

牡蛎礁是维持生物多样性、净化水质和保护海岸线的关键生态系统,但全球范围正持续退化。恢复此类生态需定期进行水下监测,但人工潜水成本高、风险大且覆盖有限。自主水下航行器(AUV)提供替代方案,但现有系统依赖几何导航,无法理解场景语义。近期视觉语言模型(VLM)支持智能探索,但现有端到端系统存在三大缺陷:频繁等待推理、无法建模机器人动力学导致碰撞、自我修正能力弱造成路径误差累积。为此,我们提出CORAL框架,将高层语义推理与底层反应式控制解耦。VLM负责选择航点,动力学规划器执行无碰撞路径;几何验证模块校验航点并触发重规划。相比先前最先进方法,CORAL覆盖率提升14.28个百分点(相对提升17.85%),碰撞减少100%,所需VLM调用减少57%。

原文摘要 · Abstract (English)

Oyster reefs are critical ecosystem species that sustain biodiversity, filter water, and protect coastlines, yet they continue to decline globally. Restoring these ecosystems requires regular underwater monitoring to assess reef health, a task that remains costly, hazardous, and limited when performed by human divers. Autonomous underwater vehicles (AUVs) offer a promising alternative, but existing AUVs rely on geometry-based navigation that cannot interpret scene semantics. Recent vision-language models (VLMs) enable semantic reasoning for intelligent exploration, but existing VLM-driven systems adopt an end-to-end paradigm, introducing three key limitations. First, these systems require the VLM to generate every navigation decision, forcing frequent waits for inference. Second, VLMs cannot model robot dynamics, causing collisions in cluttered environments. Third, limited self-correction allows small deviations to accumulate into large path errors. To address these limitations, we propose CORAL, a framework that decouples high-level semantic reasoning from low-level reactive control. The VLM provides high-level exploration guidance by selecting waypoints, while a dynamics-based planner handles low-level collision-free execution. A geometric verification module validates waypoints and triggers replanning when needed. Compared with the previous state-of-the-art, CORAL improves coverage by 14.28% percentage points, or 17.85% relatively, reduces collisions by 100%, and requires 57% fewer VLM calls.

水下监测视觉语言模型路径规划自主导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。