arXiv:2608.24603cs.RO2026-08

让机器人学会根据夹爪类型调整抓取策略,提升泛化能力。

Gripper-aware Vision Language Action Models

论文配图:Gripper-aware Vision Language Action Models
图 1 · 摘自论文原文
  • 引入多夹爪编码器与适配器路由,区分不同夹爪的抓取策略。
  • 在10.3万次示范数据上训练,实现跨夹爪零样本泛化。
  • 适合需要多类型夹爪适配的工业机器人场景。

视觉语言动作模型(VLAs)通过理解视觉观测与自然语言指令生成可执行动作序列,推动了通用机器人抓取与操作的发展。然而,现有VLAs通常隐含夹爪不变性假设,而实际抓取策略高度依赖于具体夹爪形态。不同夹爪类型(如平行钳、吸盘)需采用不同交互策略达成相同目标。此外,当前数据集主要基于平行钳夹爪,限制了夹爪感知学习。为此,我们提出MiGA多夹爪感知数据集,覆盖五种不同夹爪类型,涵盖多个机器人,包含103,000次示范,明确捕捉相同任务目标下的策略差异。我们进一步提出GVLA,结合新型多夹爪分词器与基于适配器的动作路由机制。新夹爪编码生成结构化嵌入,平衡参数共享与策略分化,层间探测验证了有意义的夹爪条件表征。大量仿真与真实机器人实验表明,GVLA在所有评估设置中均优于现有基线,且显著提升对新物体或未见任务的零样本泛化与少样本适应能力,并支持更高效的夹爪切换。

原文摘要 · Abstract (English)

Vision language action models (VLAs) have advanced general purpose robotic grasping and manipulation by enabling robots to interpret visual observations and natural language instructions to generate executable action sequences. However, existing VLAs often implicitly assume gripper invariance, despite grasping strategies being inherently embodiment-dependent. Different gripper types, such as parallel-jaw and suction, usually require distinct interaction strategies to achieve the same grasping objective. Moreover, current datasets for VLAs predominantly rely on parallel-jaw grippers, limiting gripper-aware learning. To address this gap, we introduce MiGA, a multi-gripper-aware dataset spanning five distinct gripper types across multiple robots with 103,000 demonstrations, explicitly capturing strategy divergence under shared task objectives. We further propose GVLA, which combines a new multi-gripper tokenizer with adapter-based policy routing. Our new gripper encoding induces structured embedding information that balances parameter sharing and strategy differentiation, while layer-wise probing confirms meaningful gripper-conditioned representations for VLAs. Intensive experiments in both simulation and real-world robots show that our GVLA outperforms the current baselines across evaluated settings. Our method also improves zero-shot generalization or few-shot adaptation to new objects or unseen tasks, and enable more efficient gripper adaptation.

机器人抓取多模态具身智能零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。