arXiv:2603.27012cs.ROcs.AI2026-03被引 2

用陆上演示数据训练水下抓取模型,无需真实水下操作

UMI-Underwater: Learning Underwater Manipulation without Underwater Teleoperation

  • 自监督收集水下成功抓取数据,减少真实场景采集成本
  • 通过深度感知表征实现陆地到水下的零样本迁移,性能优于纯RGB基线
  • 适合缺乏水下实操数据的研究者和工业应用开发者

水下机器人抓取因图像质量差、变化大以及示范数据获取成本高而困难。本文提出一个系统:(i) 通过自监督数据收集管道自动获取成功的水下抓取示范;(ii) 利用基于深度的可操作性表征,将陆上人类示范的知识迁移到水下,该表征对光照与色彩漂移具有鲁棒性。在陆地上训练的可操作性模型通过几何对齐实现水下零样本部署,随后基于水下示范训练条件扩散策略以生成控制动作。池中实验表明,该方法提升抓取性能与背景变化鲁棒性,并能泛化至仅在陆地数据中出现的物体,优于仅依赖RGB的基线模型。代码、视频及附加结果见 https://umi-under-water.github.io。

原文摘要 · Abstract (English)

Underwater robotic grasping is difficult due to degraded, highly variable imagery and the expense of collecting diverse underwater demonstrations. We introduce a system that (i) autonomously collects successful underwater grasp demonstrations via a self-supervised data collection pipeline and (ii) transfers grasp knowledge from on-land human demonstrations through a depth-based affordance representation that bridges the on-land-to-underwater domain gap and is robust to lighting and color shift. An affordance model trained on on-land handheld demonstrations is deployed underwater zero-shot via geometric alignment, and an affordance-conditioned diffusion policy is then trained on underwater demonstrations to generate control actions. In pool experiments, our approach improves grasping performance and robustness to background shifts, and enables generalization to objects seen only in on-land data, outperforming RGB-only baselines. Code, videos, and additional results are available at https://umi-under-water.github.io.

水下机器人零样本迁移扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。