提出双脑架构,让水下机器人更智能、更自适应地导航。
UnderwaterVLA: Dual-brain Vision-Language-Action architecture for Autonomous Underwater Navigation
- 采用双脑设计,分层处理任务推理与实时控制。
- 在浊水中导航误差降低,任务完成率提升19%至27%。
- 无需特定训练数据,适合复杂多变的水下环境。
本文提出UnderwaterVLA,一种面向自主水下导航的新框架,融合多模态基础模型与具身智能系统。水下作业因流体扰动、通信带宽有限及浑浊水域传感退化而困难重重。为此,提出三项创新:首先,双脑架构将高层任务推理与底层反应式控制解耦,提升在通信与计算受限下的鲁棒性;其次,首次将视觉-语言-动作(VLA)模型应用于水下机器人,引入结构化思维链推理,实现可解释决策;第三,设计受流体力学启发的模型预测控制(MPC)方案,在不需昂贵任务特训的前提下实时补偿流体影响。实地测试结果表明,UnderwaterVLA在视觉条件劣化时降低导航误差,同时任务完成率较基线提升19%至27%。通过减少对水下专用训练数据的依赖并增强跨环境适应性,为下一代智能无人潜水器提供可扩展、低成本的解决方案。
原文摘要 · Abstract (English)
This paper presents UnderwaterVLA, a novel framework for autonomous underwater navigation that integrates multimodal foundation models with embodied intelligence systems. Underwater operations remain difficult due to hydrodynamic disturbances, limited communication bandwidth, and degraded sensing in turbid waters. To address these challenges, we introduce three innovations. First, a dual-brain architecture decouples high-level mission reasoning from low-level reactive control, enabling robust operation under communication and computational constraints. Second, we apply Vision-Language-Action(VLA) models to underwater robotics for the first time, incorporating structured chain-of-thought reasoning for interpretable decision-making. Third, a hydrodynamics-informed Model Predictive Control(MPC) scheme compensates for fluid effects in real time without costly task-specific training. Experimental results in field tests show that UnderwaterVLA reduces navigation errors in degraded visual conditions while maintaining higher task completion by 19% to 27% over baseline. By minimizing reliance on underwater-specific training data and improving adaptability across environments, UnderwaterVLA provides a scalable and cost-effective path toward the next generation of intelligent AUVs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。