让机器人更懂3D空间,通过预测运动轨迹和环境几何提升精准操作能力。
GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation
- 引入运动轨迹与3D几何预测模块,增强模型对空间关系的感知。
- 在复杂任务中表现优于主流基线,成功率显著提升。
- 仅需轻量级查询令牌,推理高效适合真实场景部署。
视觉-语言-动作(VLA)模型在机器人操作中展现出强大泛化能力,但仍以反应式和2D为中心,难以应对需要精确3D推理的任务。本文提出GeoPredict,一种面向几何感知的VLA框架,通过引入预测性运动学与几何先验来增强连续动作策略。该框架包含两个关键模块:轨迹级模块用于编码运动历史并预测机器人臂的多步3D关键点轨迹;以及预测性3D高斯几何模块,沿未来关键点轨迹进行轨迹引导的精细建模,以预测工作空间几何结构。这些预测模块仅在训练阶段提供基于深度图渲染的监督信号,推理时无需3D解码,仅需添加轻量级查询令牌。在RoboCasa Human-50、LIBERO及真实世界操作任务上的实验表明,GeoPredict持续优于强基线模型,尤其在几何密集型与空间要求高的场景中表现突出。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models achieve strong generalization in robotic manipulation but remain largely reactive and 2D-centric, making them unreliable in tasks that require precise 3D reasoning. We propose GeoPredict, a geometry-aware VLA framework that augments a continuous-action policy with predictive kinematic and geometric priors. GeoPredict introduces a trajectory-level module that encodes motion history and predicts multi-step 3D keypoint trajectories of robot arms, and a predictive 3D Gaussian geometry module that forecasts workspace geometry with track-guided refinement along future keypoint trajectories. These predictive modules serve exclusively as training-time supervision through depth-based rendering, while inference requires only lightweight additional query tokens without invoking any 3D decoding. Experiments on RoboCasa Human-50, LIBERO, and real-world manipulation tasks show that GeoPredict consistently outperforms strong VLA baselines, especially in geometry-intensive and spatially demanding scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。