用点云补全提升3D动态模型精度,支持长时序规划与真实世界迁移。
3D Point World Models: Point Completion Enables More Accurate Dynamics Learning

- 先补全部分点云,再在完整3D场景中学习动作条件动态
- 长时序推演可达100-300步,几何一致性显著提升
- 适用于不同机器人形态,支持模拟到现实的迁移
学习世界预测模型可实现基于规划的机器人控制,使机器人能在新任务中自主设计解决方案。然而,基于视频的大规模动态模型缺乏显式的3D空间结构,在长时间推演中会出现几何不一致和误差累积。基于部分点云的新兴3D动态模型虽提升了几何一致性,但仍受遮挡影响并存在预测漂移。为此,我们提出3D Point World Models(3DPWM)——一种完全在3D空间中运行的任务无关世界模型:首先完成部分点云,然后在补全后的3D场景中学习动作条件动态。通过在完整几何结构上建模,3DPWM实现了可靠的长时序推演,并支持更准确的成本评估以用于基于模型的规划,同时具备对新任务的适应能力。在多种机器人形态和桌面操作基准上的实验表明,3DPWM能实现显著更可靠的长时序推演(100-300+步),支持开环与闭环规划,并成功实现模拟到现实的迁移。
原文摘要 · Abstract (English)
Learning predictive models of the world enables robotic control through planning, potentially allowing robots to improvise solutions on new tasks. However, large video-based dynamics models lack explicit 3D spatial structure and suffer from geometrically inconsistent long-term rollouts with compounding errors. Emerging 3D dynamics models based on partial point clouds improve geometric consistency but remain sensitive to occlusions and accumulated prediction drift. To address these challenges, we present 3D Point World Models (3DPWM) - a task-agnostic world model that operates entirely in 3D space by first completing partial point clouds and then learning action-conditioned dynamics in this completed 3D scene. By operating on completed geometry, 3DPWM enables reliable long-horizon rollouts and more accurate cost evaluation for model-based planning while supporting adaptation to new tasks. Experiments across different robotic embodiments and tabletop manipulation benchmarks demonstrate that 3DPWM achieves significantly more reliable long-horizon rollouts (100-300+ steps), supports both open-loop and closed-loop planning, and enables successful sim-to-real transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。