arXiv:2606.17046cs.ROcs.CV2026-06被引 3

用3D几何模型让机器人更懂指令和物理交互

Geometric Action Model for Robot Policy Learning

论文配图:Geometric Action Model for Robot Policy Learning
图 1 · 摘自论文原文
  • 用预训练几何模型做感知、预测和动作解码的统一框架
  • 在仿真和真实机器人任务中精度更高、速度更快、体积更小
  • 适合需要精准抓取和复杂交互的机器人控制场景

通用机器人策略需在理解用户指令的同时,推理物体、摄像头与机器人动作在三维物理世界中的交互。现有视觉-语言-动作模型(VLAs)和视频世界-动作模型(WAMs)虽继承了大规模基础模型的语义或时序先验,但仍主要在2D图像帧或2D隐空间中操作,未能显式建模接触密集型操作所需的3D几何信息。我们提出几何动作模型(GAM),一种语言条件的操控策略,直接复用预训练的几何基础模型(GFM)作为感知、时间预测和动作解码的共享基础。GAM在中间层分割GFM:浅层作为观察编码器,因果未来预测器插入该层,基于语言、本体感觉和动作历史预测未来隐状态。预测的未来隐状态通过剩余的GFM块进行特征传播与解码,使单一主干同时生成未来几何和动作。该设计通过最小架构修改赋予GFM语言条件的时间世界建模能力,同时保留其丰富的几何先验。在广泛模拟与真实机器人操控基准测试中,GAM比当前大模型规模基线更准确、更鲁棒、更快、更轻量。

原文摘要 · Abstract (English)

Generalist robot policies must follow user instructions while reasoning about how objects, cameras, and robot actions interact in the 3D physical world. Recent vision-language-action models (VLAs) and video world-action models (WAMs) inherit strong semantic or temporal priors from large-scale foundation models, but they still operate primarily on 2D image frames or 2D-derived latent spaces, leaving implicit the 3D geometry required for contact-rich manipulation. We propose the Geometric Action Model (GAM), a language-conditioned manipulation policy that directly repurposes a pretrained geometric foundation model (GFM) as a shared substrate for perception, temporal prediction, and action decoding. GAM splits the GFM at an intermediate layer: the shallow layers serve as an observation encoder, and a causal future predictor inserted at the split layer forecasts future latent tokens conditioned on language, proprioception, and action history. The predicted future tokens are then routed through the remaining GFM blocks for feature propagation and decoding, allowing a single backbone to produce both future geometry and actions. This design equips the GFM with language-conditioned temporal world modeling through minimal architectural modification while preserving its rich geometric priors. Across a broad suite of simulation and real-robot manipulation benchmarks, GAM is more accurate, more robust, faster, and lighter than current foundation-model-scale baselines.

机器人学习3D几何语言控制动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。