让机器人模型无视自身硬件,跨设备零样本操控物体
Cloak: Zero-Shot Cross-Embodiment Manipulation by Masking the End-Effector from the VLA

- 用模拟遮罩隐藏机械臂末端,实现视觉与身体解耦
- 仅用一个夹爪数据集训练,即可零样本迁移到多种新硬件
- 适合需要快速适配新机器人、避免重复采集数据的场景
我们提出Cloak,一种训练方法,使视觉-语言-动作(VLA)模型在不依赖特定硬件的情况下,实现零样本跨体感迁移。该方法通过在腕部摄像头视角中掩蔽机械臂末端,消除其对视觉输入的影响。末端执行器在腕视图中占据大而稳定的区域,使用已知机器人几何结构实时生成精确遮罩,无需分割或生成模型。训练时引入遮罩增强,提升模型对未见体感的泛化能力。我们基于单一平行夹爪数据集训练了Cloak-VLA,未收集任何新体感数据。结果表明,Cloak-VLA可零样本迁移到另一夹爪、另一机械臂及五指手等未见体感,且保留原体感性能。通过解耦腕视图与本体,使数据脱离硬件限制,可长期复用。
原文摘要 · Abstract (English)
We present Cloak, a training recipe that endows a Vision-Language-Action (VLA) model with zero-shot cross-embodiment transfer by cloaking the end-effector from its own wrist camera. The end-effector occupies a large and consistent region of the wrist view and masking it allows for embodiment-agnostic visual reasoning. Cloak renders a mask in simulation from the robot's known geometry, accurately and in real time, with no segmentation or generative models. During training, we augment the mask so the model generalizes to embodiments unseen at training time. We demonstrate the recipe with Cloak-VLA, a VLA trained with Cloak on a single parallel-jaw gripper dataset. No data of new embodiments is ever collected. Cloak-VLA transfers zero-shot to various unseen embodiments, including another gripper, another arm, and a five-fingered hand, while preserving the source embodiment's performance. By decoupling the wrist view from its own embodiment, Cloak allows data to outlive the hardware it was collected on.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。