arXiv:2607.07101cs.ROcs.AI2026-07

让机器人视觉与自身状态对齐,提升操作精度。

GeoProp: Grounding Robot State in Vision for Generalist Manipulation

论文配图:GeoProp: Grounding Robot State in Vision for Generalist Manipulation
图 1 · 摘自论文原文
  • 通过几何投影将机器人状态映射到图像,生成有位置依据的视觉特征。
  • 在仿真和真实场景中,使扩散策略性能平均提升10.6%。
  • 轻量级模块,仅增加2-3%参数,适合通用机器人控制任务。

本研究提出GeoProp,一种轻量级、可即插即用的适配器,用于将机器人本体感知(proprioception)与视觉信息显式对齐。传统融合方法将本体感知视为孤立向量,缺乏与视觉特征图的直接对应关系,导致操控策略难以在场景中定位机器人状态,性能常低于纯视觉基线。GeoProp通过将机器人状态投影至图像平面,采样局部视觉特征并构建具有空间锚定的“状态令牌”;再利用FiLM调制将状态相关的空间先验注入视觉特征。为捕捉运动意图,还基于近期运动学预测短时前瞻坐标,提供预判性视觉上下文。在67项任务中,该方法使扩散策略在63个仿真任务上提升8.7%,在RoboTwin子集上使pi_0提升4.0%,真实世界平均增益达10.6%,且仅增加2-3%模型参数。结果表明,这种几何接地机制是通用具身策略中简单而高效的归纳偏置。

原文摘要 · Abstract (English)

Proprioception is fundamental to robotic manipulation, yet standard fusion methods often treat it as an isolated vector lacking explicit alignment with visual tokens. Without a direct correspondence between 3D kinematics and 2D feature maps, manipulation policies struggle to ground the robot's state within the scene, frequently underperforming even vision-only baselines. To address this, we introduce GeoProp, a lightweight, plug-and-play adapter that aligns proprioception with vision through explicit geometric grounding and spatial feature sampling. GeoProp projects the robot state onto the image plane to sample localized visual features, constructing a grounded state token. It then injects state-derived spatial priors into the corresponding visual features via FiLM modulation. To capture motion intent, GeoProp further samples features at a short-horizon predicted coordinate derived from recent kinematics, providing look-ahead visual context. Across 67 tasks, GeoProp improves Diffusion Policy by 8.7% on 63 simulation tasks and pi_0 by 4.0% on the RoboTwin subset, and yields a 10.6% average gain across both policy families in the real world, while adding only 2-3% to the parameter count. These results demonstrate that GeoProp is a simple yet high-impact inductive bias for generalist embodied policies. Project page: https://alibaba-damo-academy.github.io/GeoProp/.

机器人操作视觉对齐扩散策略本体感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。