用3D几何表示替代语言/视频模型,让机器人更精准地从视觉直接生成动作。
Robotic Manipulation is Vision-to-Geometry Mapping: Vision-Geometry Backbones over Language and Video Models
- 用预训练的3D世界模型替代传统语言或视频骨干网络。
- 在仿真和真实场景中均超越多个领先基线模型,尤其在未见视角下表现优异。
- 适合追求高精度物理操作与通用机器人控制的研究者。
机器人操作本质上是视觉到几何的映射($f(v) \rightarrow G$)。物理动作的根本属性是3D位置与空间关系等几何特征。因此我们主张,通用机器人控制的基础应是视觉-几何骨干网络,而非广泛采用的视觉-语言或视频模型。传统VLA和视频预测模型依赖大规模2D图像-文本或时序像素数据预训练,其表征受语义概念或2D先验主导,难以契合物理操作所需的精确3D几何特性。为此,我们提出视觉-几何-动作(VGA)模型,直接以预训练3D表征作为动作生成条件。具体而言,VGA将传统语言或视频骨干替换为预训练3D世界模型,实现视觉输入到物理动作的无缝映射。为进一步增强几何一致性,引入渐进式体素调制,并联合训练动作与3D属性预测以保持几何表征。大量实验验证了该方法的有效性:在仿真基准上,VGA优于π_{0.5}、OpenVLA-OFT、GeoVLA和Motus等主流基线;在真实部署中,其在未见视角下超越π_{0.5},并能准确执行目标抓取的语言指令。结果表明,基于原生3D表示而非语言或视频先验,是迈向可泛化的物理智能的重要方向。
原文摘要 · Abstract (English)
At its core, robotic manipulation is a problem of vision-to-geometry mapping ($f(v) \rightarrow G$). Physical actions are fundamentally defined by geometric properties like 3D positions and spatial relationships. Consequently, we argue that the foundation for generalizable robotic control should be a vision-geometry backbone, rather than the widely adopted vision-language or video models. Conventional VLA and video-predictive models rely on backbones pretrained on large-scale 2D image-text or temporal pixel data. While effective, their representations are largely shaped by semantic concepts or 2D priors, which do not intrinsically align with the precise 3D geometric nature required for physical manipulation. Driven by this insight, we propose the Vision-Geometry-Action (VGA) model, which directly conditions action generation on pretrained 3D representations. Specifically, VGA replaces conventional language or video backbones with a pretrained 3D world model, establishing a seamless vision-to-geometry mapping that translates visual inputs directly into physical actions. To further enhance geometric consistency, we introduce Progressive Volumetric Modulation and jointly train action and 3D property prediction to preserve geometric representations. Extensive experiments validate the effectiveness of our approach. Across simulation benchmarks, VGA outperforms leading VLA, 3D-VLA, and WAM baselines, including $π_{0.5}$, OpenVLA-OFT, GeoVLA, and Motus. In real-world deployments, VGA surpasses $π_{0.5}$ under unseen viewpoints and accurately follows language instructions for target grasping. These results highlight that operating on native 3D representations, rather than relying primarily on language or video priors, offers a promising direction toward generalizable physical intelligence. Project page: https://hcplab-sysu.github.io/VisionGeometryActionModel.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。