用3D点云动态实现视频到机器人动作的精准映射。
PointAction: 3D Points as Universal Action Representations for Robot Control

- 将视频预测与3D点云动态结合,生成带时空一致性的4D场景变化。
- 在仿真和真实机器人上均实现领先性能,支持跨任务、跨机械臂泛化。
- 仅需少量动作标注即可迁移,适合多场景机器人控制应用。
视频动作模型(VAMs)利用预训练视频扩散模型捕捉的广泛视觉动态,为通用机器人操作提供了有前景的路径。然而,仅基于RGB的视频回放无法直接执行:它们未能明确指定度量3D运动、接触几何和精细空间约束,导致动作定位模糊。同时,跨多样化任务和机器人形态扩展动作监督仍成本高昂。我们提出PointAction,通过显式的点基4D建模,连接视频预测与机器人动作。PointAction微调基础视频生成模型,联合预测未来RGB帧与动态3D点图,生成任务相关场景几何的时序一致3D运动。这些点动态作为结构化、形态无关的动作接口,由基于扩散的动作解码器映射为可执行机器人动作。通过使用度量3D点动态作为视频预测与控制之间的接口,PointAction降低了仅基于RGB动作定位的模糊性,并在有限动作监督下支持跨任务与跨形态的迁移。实验表明,PointAction在机器人场景中实现了最先进的4D生成质量,在仿真中优于现有基线,并泛化至两个预训练中未见的真实机械臂。
原文摘要 · Abstract (English)
Video-Action Models (VAMs) leverage the broad visual dynamics captured by pre-trained video diffusion models, offering a promising path toward generalizable robot manipulation. However, RGB-only video rollouts are not directly actionable: they leave metric 3D motion, contact geometry, and fine-grained spatial constraints under-specified, making action grounding ambiguous. Meanwhile, scaling action supervision across diverse tasks and embodiments remains costly. We present PointAction, a framework that bridges video predictions to robot actions through explicit point-based 4D modeling. PointAction fine-tunes a foundation video generation model to jointly predict future RGB frames and dynamic 3D pointmaps, producing temporally consistent 3D motion of task-relevant scene geometry. These point dynamics serve as a structured, embodiment-agnostic action interface, which a diffusion-based action decoder maps to executable robot actions. By using metric 3D point dynamics as the interface between video prediction and control, PointAction reduces the ambiguity of RGB-only action grounding and supports transfer across tasks and embodiments with limited action supervision. Experiments show that PointAction achieves state-of-the-art 4D generation quality on robot scenes, outperforms existing baselines in simulation, and generalizes to two real robot arms unseen during pretraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。