arXiv:2607.04714cs.ROcs.AI2026-07

通过预测点云变化学习机器人抓取的运动隐变量,提升泛化能力。

Geometry-Aware Motion Latents for Learning Robust Manipulation Policies

论文配图:Geometry-Aware Motion Latents for Learning Robust Manipulation Policies
图 1 · 摘自论文原文
  • 用点云演化预测替代视觉重建,学习三维几何变化的离散运动隐变量
  • 仅用单视角RGB-D图像达到顶尖性能,多任务测试成功率超90%
  • 适合需要少样本、强鲁棒性的真实场景机械臂控制研究者

机器人抓取动作的学习通常依赖从视觉序列中提取运动模式,但有效的动作抽象需理解三维几何变换。本文提出GeoMoLa(Geometry-Aware Motion Latents),通过预测点云在操作过程中的演变来学习离散的运动隐变量,而非重建视觉观测。这一四维目标——随时间变化的空间几何——迫使隐表示编码真实的物理运动,而非外观模式。GeoMoLa仅使用单视角RGB-D输入即达到当前最优性能,而现有方法需多视角重建;在多种抓取基准上表现优异。消融实验表明,几何预测是性能提升的关键,定量验证了抓取任务依赖空间理解。所学隐变量具备良好运动抽象能力:在新场景中应用仍产生符合物理规律的变换,与视觉上下文无关。真实世界实验进一步证实其鲁棒性,在杂乱环境中仅需少量示范即可成功执行任务。结果表明,通过理解运动的三维效应而非像素级模式,可更有效地生成用于机器人控制的运动隐变量。

原文摘要 · Abstract (English)

Learning motion latents for robotic manipulation heavily relies on extracting motion patterns from visual sequences, yet effective action abstractions require understanding three-dimensional geometric transformations. Here, we introduce GeoMoLa (Geometry-Aware Motion Latents), which learns discrete motion latent codes by predicting how point clouds evolve during manipulation rather than reconstructing visual observations. This four-dimensional objective -- spatial geometry changing through time -- forces latent representations to encode actual physical motion rather than appearance patterns. GeoMoLa achieves state-of-the-art performance using only single-view RGB-D input, while existing methods require multi-view reconstruction, succeeding across diverse manipulation benchmarks. Our ablations reveal that geometric prediction is the key to driving performance, quantitatively validating that manipulation depends on spatial understanding. Furthermore, the learned codes exhibit effective motion abstraction: applying them to novel scenes produces physically consistent transformations regardless of visual context. Our real-world experiments also confirm this robustness capability, achieving robust manipulation with minimal demonstrations in cluttered environments where geometric reasoning determines success. Thus, we demonstrate that effective motion latents for robot control can better emerge from understanding motion through its three-dimensional effects rather than pixel-level patterns.

机器人控制运动隐变量几何感知少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。