arXiv:2609.02546cs.RO2026-09

对比两种零样本迁移方式,提升机器人抓取任务的泛化能力

ZETA: A Controlled Study of Zero-Shot Cross-Embodiment VLA Transfer for Tabletop Manipulation

论文配图:ZETA: A Controlled Study of Zero-Shot Cross-Embodiment VLA Transfer for Tabletop Manipulation
图 1 · 摘自论文原文
  • 区分严格与预训练暴露的零样本迁移,明确评估标准
  • 局部末端执行器状态表示等方法提升性能15~18个百分点
  • 仅用5%目标机器人数据预训练,进度提升13.4个百分点

零样本泛化到未见机器人本体对通用视觉-语言-动作模型至关重要,因硬件迭代频繁且任务数据采集成本高。然而,该问题缺乏统一定义与受控评估环境,常混淆本体差异与其他变量。本文首先区分严格零样本(目标本体完全未出现在训练中)与预训练暴露零样本(仅在预训练中出现)。构建覆盖14个未见本体的模拟与真实世界评估基准,在此框架下分析四类因素:状态-动作表示、预训练本体多样性、辅助协同训练目标、目标本体曝光程度。实验表明,局部末端执行器状态表示、源本体多样性及辅助协同训练分别提升跨本体迁移效果约15、18和7个百分点;仅在预训练中加入5%目标本体数据,平均进度提升13.4个百分点,证明两类零样本迁移本质不同,应分别报告。研究为双指夹持器在固定桌面上的跨本体迁移提供实用指导,并推动对移动基座、灵巧手及长时序任务等更广场景的探索。

原文摘要 · Abstract (English)

Zero-shot generalization to unseen embodiments is important for generalizable vision-language-action (VLA) models as robot hardware evolves and task-specific data collection remains costly. However, a systematic understanding of this problem remains limited, in part because the literature lacks a unified zero-shot transfer definition and controlled evaluation settings that isolate embodiment changes from differences in tasks, scenes, or protocols. To address this gap, we first distinguish strict zero-shot transfer, where the target embodiment is absent from all training data, from pretrain-exposed zero-shot transfer, where it appears only during pretraining. We then introduce a controlled benchmark spanning 14 held-out target embodiments across simulation and real-world validation. Within this framework, we conduct a controlled analysis of four factors: state-action representations, pretraining embodiment diversity, auxiliary co-training objectives, and target-embodiment exposure. Experimental results show that local end-effector (EEF) state-action representations, the source embodiment diversity, and auxiliary co-training improve cross-embodiment transfer by around 15, 18, and 7 percentage points, respectively. We further find that adding only 5% target-embodiment data during pretraining improves average target-embodiment progress by 13.4 percentage points, showing that strict and pretrain-exposed zero-shot transfer are distinct and should be reported separately. Together, these findings provide practical guidance for evaluating and improving cross-embodiment VLA transfer in stationary tabletop manipulation with two-finger grippers, while motivating future investigation of broader settings including mobile-base control, dexterous hands, and long-horizon tasks.

零样本迁移机器人控制视觉语言动作泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。