arXiv:2606.24472cs.ROcs.AI2026-06被引 1

让视觉语言动作模型理解相机几何结构,提升多视角机器人操作精度。

G$^3$VLA: Geometric inductive bias for Vision-Language-Action Models

论文配图:G$^3$VLA: Geometric inductive bias for Vision-Language-Action Models
图 1 · 摘自论文原文
  • 引入相机内参外参信息,将2D图像特征映射到真实空间坐标。
  • 在多个机器人任务中实现稳定提升,尤其在空间敏感任务上效果显著。
  • 无需深度传感器或人工标注,适用于真实机器人部署场景。

视觉-语言-动作(VLA)模型通过预训练的视觉-语言骨干网络快速推进通用机器人操作,但其视觉标记仍基于二维图像坐标,而非机器人相机的校准几何结构——这一问题在多相机设置中尤为突出,尽管各视角间存在已知的内参与外参关系,却仍被当作独立图像处理。本文提出G$^3$VLA,一种相机感知的几何模块,可在不改变动作空间或模仿目标的前提下,将校准结构注入预训练VLA的视觉标记流中,结合内参条件化的射线嵌入、投影位置编码(PRoPE)和双向跨视图融合。几何监督可来自真实点图(若有),或通过置信度门控的$π^3$X教师预测,无需深度传感器或人工标注。在$π_0$上实例化后,G$^3$VLA在LIBERO套件、RoboCasa24、RoboTwin2.0及真实机器人设置中均取得一致性能提升,空间敏感任务改善最为明显。进一步在$π_{0.5}$和GR00T 1.5上验证,结果表明当几何感知标记能直接接入动作生成路径时,几何迁移效果最佳。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models have made rapid progress in generalist robot manipulation by harnessing semantic knowledge from pretrained vision-language backbones, but their visual tokens remain grounded in 2D image coordinates rather than the calibrated geometry of the robot's cameras -- a mismatch especially pronounced in multi-camera setups, where views are coupled by known intrinsics and extrinsics yet processed as independent images. We propose G$^3$VLA, a camera-aware geometric module that injects calibrated structure into the visual-token stream of a pretrained VLA without altering its action space or imitation objective, combining intrinsic-conditioned ray embeddings, projective positional encoding (PRoPE), and bidirectional cross-view fusion. Geometric supervision is provided either from ground-truth point maps when available, or from confidence-gated $π^3$X teacher predictions, requiring no depth sensors or manual annotations. Instantiated on $π_0$, G$^3$VLA yields consistent gains across the LIBERO suites, RoboCasa24, RoboTwin2.0, and real-robot settings, with the largest improvements on spatially and object-sensitive tasks. We further validate on $π_{0.5}$ and GR00T 1.5, with results suggesting that geometric transfer is most effective when geometry-aware tokens have direct access to the action generation pathway. Our project page is at https://sites.google.com/view/g3vla

机器人操作多视角学习几何先验视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。