arXiv:2509.13666cs.ROcs.AI2025-09被引 4

用视觉语言模型让水下机器人智能找目标,省时省力还更准。

DREAM: Domain-aware Reasoning for Efficient Autonomous Underwater Monitoring

  • 用VLM驱动机器人自主决策,不依赖预先定位信息。
  • 找牡蛎任务节省31.5%时间,覆盖更多目标且步数减少23%。
  • 在沉船场景中实现100%覆盖率,远超传统方法的60.23%。

海洋正在变暖和酸化,对温度敏感的贝类(如牡蛎)构成严重威胁,亟需长期监测。但人工潜水成本高、风险大,因此采用机器人方案更安全高效。为使水下机器人能实时做出环境感知决策,需赋予其智能‘大脑’,以实现持久、广域、低成本的海底监测。为此,我们提出DREAM框架——一种基于视觉语言模型(VLM)的自主导航系统,用于长期水下探索与栖息地监测。实验表明,该框架在无先验位置信息的情况下仍能高效发现并探索目标(如牡蛎、沉船)。在牡蛎监测任务中,相比基线方案,耗时减少31.5%,且同等采样量下覆盖更多目标;相较原生VLM,步骤减少23%,目标覆盖提升8.88%。在沉船场景中,系统成功完成无碰撞探索与建图,步数减少27.5%,覆盖率达100%;而原生模型平均仅达60.23%。

原文摘要 · Abstract (English)

The ocean is warming and acidifying, increasing the risk of mass mortality events for temperature-sensitive shellfish such as oysters. This motivates the development of long-term monitoring systems. However, human labor is costly and long-duration underwater work is highly hazardous, thus favoring robotic solutions as a safer and more efficient option. To enable underwater robots to make real-time, environment-aware decisions without human intervention, we must equip them with an intelligent "brain." This highlights the need for persistent,wide-area, and low-cost benthic monitoring. To this end, we present DREAM, a Vision Language Model (VLM)-guided autonomy framework for long-term underwater exploration and habitat monitoring. The results show that our framework is highly efficient in finding and exploring target objects (e.g., oysters, shipwrecks) without prior location information. In the oyster-monitoring task, our framework takes 31.5% less time than the previous baseline with the same amount of oysters. Compared to the vanilla VLM, it uses 23% fewer steps while covering 8.88% more oysters. In shipwreck scenes, our framework successfully explores and maps the wreck without collisions, requiring 27.5% fewer steps than the vanilla model and achieving 100% coverage, while the vanilla model achieves 60.23% average coverage in our shipwreck environments.

水下机器人视觉语言模型自主导航环境监测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。