arXiv:2510.07869cs.RO2025-10被引 4

构建水下机器人通用智能数据集与模型,支持语言指令下的多任务执行。

USIM and U0: A Vision-Language-Action Dataset and Model for General Underwater Robots

  • 提出USIM数据集与U0视觉-语言-动作模型,融合语言指令驱动感知与操作。
  • 在水下任务中实现43.1%在线成功率,导航任务达87.5%,较基线提升5.5%。
  • 引入目标位姿估计辅助任务,增强模型空间理解能力,适合水下机器人研究者。

水下环境对机器人导航与操作带来独特挑战。现有研究多聚焦特定任务,缺乏通用智能的多任务执行研究。为此,我们提出统一框架,集成感知与行动以响应语言指令。首先,构建基于仿真、包含2275条轨迹、超过905,000帧、总计约25小时BlueROV2交互的USIM数据集。其次,提出U0视觉-语言-动作(VLA)模型,可执行避障导航至三维移动操作等多种任务。该模型采用卷积-注意力感知模块(CAP),通过目标位姿估计作为辅助任务,显式增强空间意识。评估方面,建立涵盖离线指标与在线任务执行的系统化评测框架与自动化流程。实验表明,USIM显著提升现有VLA模型在水下场景的适应能力。尤其,U0模型将离线动作预测误差降至0.0359,整体在线成功率达43.1%,较基线(<37.6%)提升5.5%,导航任务最高达87.5%。结果验证了水下机器人通用智能的可行性,为可扩展数据合成与水生具身智能体奠定基础。

原文摘要 · Abstract (English)

Underwater environments pose unique challenges for robotic navigation and manipulation. While existing research has primarily focused on task-specific methods, studies on general-purpose intelligence for multi-task execution remain scarce. To address this gap, we propose a unified framework for general-purpose underwater robots that integrates perception and action driven by language instructions. First, we develop a data synthesis pipeline to construct USIM, a simulation-based dataset which comprises over 905K frames from 2275 trajectories, totaling approximately 25 hours of BlueROV2 interactions. Furthermore, we propose U0, a vision-language-action (VLA) model capable of executing various tasks from obstacle-avoidance navigation to three-dimensional mobile manipulation. The model features a convolution-attention-based perception (CAP) module, which incorporates target pose estimation as an auxiliary task to explicitly bolster the model's spatial awareness. For evaluation, we establish a systematic assessment framework and an automated pipeline encompassing both offline metrics and online task execution. Experimental results demonstrate that the USIM dataset significantly empowers existing VLA models to adapt to underwater scenarios. Notably, our U0 model achieves state-of-the-art performance: it reduces the offline mean action prediction error to 0.0359 and achieves an overall online success rate of 43.1%, marking a 5.5% improvement over existing competitive baselines (below 37.6%), with navigation tasks reaching as high as 87.5%. These results validate the feasibility of general-purpose intelligence in underwater robotics, providing a foundation for scalable dataset synthesis and aquatic embodied agents.

水下机器人视觉语言动作数据集具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。