arXiv:2601.03782cs.ROcs.AI2026-01被引 71

用点云流预测3D世界响应,让机器人从单张图就能动手操作。

PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation

  • 用3D点流统一状态与动作,直接关联物理几何
  • 训练数据达200万轨迹、500小时,覆盖真实与仿真场景
  • 单张图+零样本即可完成推拉、工具使用等复杂操作

人类仅凭一眼和身体动作预判3D世界反应,这对机器人操控同样关键。我们提出PointWorld,一个大规模预训练的3D世界模型,将状态与动作统一表示为3D点流:给定一至几张RGB-D图像及低层机器人动作指令,模型可预测每像素在3D空间中的位移响应。通过将动作表示为3D点流而非特定机械臂的动作空间(如关节位置),该方法直接依赖机器人的物理几何,同时实现跨机器人形态的无缝学习。为训练此3D世界模型,我们构建了涵盖真实与模拟环境的大型数据集,得益于近期3D视觉与仿真技术进展,数据总计约200万条轨迹,累计500小时,覆盖单臂Franka与双臂人形机器人。通过大规模实证研究骨干网络、动作表示、学习目标、部分可观测性、数据混合、域迁移与规模扩展,我们提炼出大规模3D世界建模的设计原则。具备实时推理能力(0.1秒),PointWorld可高效集成至模型预测控制(MPC)框架中进行操控。实验表明,仅用一个预训练模型,真实世界的Franka机器人即可完成刚体推移、柔性和关节物体操作以及工具使用,无需任何演示或微调,且仅需一张野外拍摄的图像。项目主页见 https://point-world.github.io/。

原文摘要 · Abstract (English)

Humans anticipate, from a glance and a contemplated action of their bodies, how the 3D world will respond, a capability that is equally vital for robotic manipulation. We introduce PointWorld, a large pre-trained 3D world model that unifies state and action in a shared 3D space as 3D point flows: given one or few RGB-D images and a sequence of low-level robot action commands, PointWorld forecasts per-pixel displacements in 3D that respond to the given actions. By representing actions as 3D point flows instead of embodiment-specific action spaces (e.g., joint positions), this formulation directly conditions on physical geometries of robots while seamlessly integrating learning across embodiments. To train our 3D world model, we curate a large-scale dataset spanning real and simulated robotic manipulation in open-world environments, enabled by recent advances in 3D vision and simulated environments, totaling about 2M trajectories and 500 hours across a single-arm Franka and a bimanual humanoid. Through rigorous, large-scale empirical studies of backbones, action representations, learning objectives, partial observability, data mixtures, domain transfers, and scaling, we distill design principles for large-scale 3D world modeling. With a real-time (0.1s) inference speed, PointWorld can be efficiently integrated in the model-predictive control (MPC) framework for manipulation. We demonstrate that a single pre-trained checkpoint enables a real-world Franka robot to perform rigid-body pushing, deformable and articulated object manipulation, and tool use, without requiring any demonstrations or post-training and all from a single image captured in-the-wild. Project website at https://point-world.github.io/.

3D世界模型机器人操控点云流零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。