arXiv:2606.08530cs.ROcs.AI2026-06被引 2

让机器人学会理解物体空间关系,提升对陌生物体和设备的操控通用性。

GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation

论文配图:GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation
图 1 · 摘自论文原文
  • 通过分步学习动作语义与连续轨迹,构建统一的空间感知动作表示。
  • 在212个未见物体上实现90.1%抓取成功率,真实机器人上达85.9%成功。
  • 适合需要跨设备、跨场景通用操控的机器人研究与应用开发者。

视觉-语言-动作(VLA)模型在基准测试中表现优异,但在真实部署中仍面临未见物体、背景变化和不同机器人形态的挑战。我们认为根源在于缺乏统一的几何感知操作表征,导致现有VLA模型易受低层轨迹监督、3D特征错位及设备差异影响。为此,我们提出GEAR-VLA框架,通过粗到细的动作学习,利用多源具身预训练赋予视觉语言模型具身推理与离散动作理解能力,再将动作语义连接至解耦梯度的DiT连续动作专家。同时,通过可训练3D空间骨干与VLA表示对齐,冻结原有视觉路径以实现语义对齐的3D融合。为实现跨机器人共享,采用形态归一化机制,将设备差异限制在底层接口,保持动作不变。大量仿真与真实实验表明其强大泛化能力:在LIBERO、零样本LIBERO-Plus和RoboTwin 2.0上达到当前最优表现,在AgileX上达85.9%成功率,在未预训练的LDT-01设备上达81.0%,并在包含6,360次试验的通用抓取基准上实现90.1%成功率(212种未见物体)。代码与模型将在https://github.com/babynabeauty/GEAR-VLA发布。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world deployment with unseen objects, background shifts, and different robot embodiments. We argue that this stems from the lack of a unified geometry-aware manipulation representation, leaving existing VLAs vulnerable to low-level trajectory supervision, misaligned 3D features, and embodiment differences. To address this, we propose GEAR-VLA, a VLA framework for learning unified geometry-aware action representations for generalizable robotic manipulation. GEAR-VLA adopts coarse-to-fine action learning, where multi-source embodied pretraining equips the VLM with embodied reasoning and discrete action understanding before latent action tokens connect action semantics to a gradient-decoupled DiT continuous action expert. It further performs semantic-aligned 3D integration by aligning a trainable 3D spatial backbone with the VLA representation while freezing the original VLM-aligned visual pathway. To share this representation across robots, GEAR-VLA uses embodiment canonicalization, where embodiment-aware states and embodiment-invariant actions confine robot differences to the low-level interface. Extensive simulation and real-world experiments demonstrate strong generalization: GEAR-VLA achieves state-of-the-art performance on LIBERO, zero-shot LIBERO-Plus, and RoboTwin 2.0, reaches 85.9% success on AgileX and 81.0% on the pretraining-unseen LDT-01 embodiment, and obtains 90.1% success on a 6,360-trial universal grasping benchmark with 212 unseen objects. Code and models will be released at https://github.com/babynabeauty/GEAR-VLA.

机器人操控几何感知通用性具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。