arXiv:2606.03240cs.RO2026-06

让视觉语言动作模型学会精准空间对齐,提升机器人操作准确性。

GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models

论文配图:GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models
图 1 · 摘自论文原文
  • 用机器人真实深度数据后训练视觉分支,生成带几何信息的特征。
  • 在真实场景任务中达78.8%成功率,显著优于仅依赖语义的模型。
  • 适合需要精确空间操作的机器人控制研究者参考。

当前视觉-语言-动作(VLA)模型多聚焦语义对齐,而可执行操作还需几何感知的空间对齐与动态可用性选择。本文提出GeoAlign,一种基于状态引导的空间对齐架构,用于VLA策略学习。该方法通过机器人领域的真实RGB-D数据对RGB分支进行后训练,生成基于RGB的几何增强后训练(GEP)特征,用于策略推理。机器人的本体感知状态查询GEP特征网格,生成紧凑且相位相关的几何令牌,用于动作预测。GeoAlign在LIBERO数据集上达到99.0%成功率,在三个SimplerEnv-Fractal任务中平均达85.3%,在八个几何敏感的真实世界ALOHA任务中达78.8%。消融实验验证了几何后训练和本体状态引导查询的有效性。

原文摘要 · Abstract (English)

Current Vision--Language--Action (VLA) models often optimize for semantic grounding, whereas executable manipulation requires geometry-aware spatial alignment and dynamic affordance selection. We introduce GeoAlign, a state-guided spatial alignment architecture for VLA policy learning. GeoAlign post-trains an RGB geometry branch with robot-domain RGB-D supervision, yielding RGB-derived Geometry-Enhanced Post-Trained (GEP) features for policy rollout. The robot's proprioceptive state queries the GEP feature grid, producing compact, phase-dependent geometry tokens for action prediction. GeoAlign achieves 99.0% on LIBERO, 85.3% across three SimplerEnv-Fractal tasks, and 78.8% on eight geometry-critical real-world ALOHA tasks, with ablations confirming the value of geometry post-training and proprioceptive-state-guided querying.

机器人控制空间对齐几何感知VLA模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。