arXiv:2508.17600cs.ROcs.AI2025-08ICCV被引 49

用高斯原语构建可扩展的3D世界模型,提升机器人操作预测与训练效果。

GWM: Towards Scalable Gaussian World Models for Robotic Manipulation

论文配图:GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
图 1 · 摘自论文原文
  • 通过高斯原语和扩散Transformer建模3D世界动态变化
  • 在模拟与真实场景中均实现精准未来状态预测
  • 适用于模仿学习与基于模型的强化学习,适合机器人控制研究者

在真实世界交互效率低下的背景下,基于图像的世界模型与策略已展现潜力,但缺乏对三维空间中几何信息的稳健理解,即使在互联网规模视频上预训练亦如此。为此,我们提出一种新型世界模型——高斯世界模型(Gaussian World Model, GWM),通过推断机器人动作下高斯原语的传播来重建未来状态。其核心为结合3D变分自编码器与潜在扩散变换器(DiT),实现基于高斯点云的细粒度场景级未来状态重建。GWM不仅能通过自监督未来预测增强视觉表征,用于模仿学习;还可作为神经模拟器支持基于模型的强化学习。模拟与真实实验表明,GWM能准确预测多种机器人动作下的未来场景,并进一步训练出性能显著超越现有方法的策略,展现出3D世界模型初步的数据扩展潜力。

原文摘要 · Abstract (English)

Training robot policies within a learned world model is trending due to the inefficiency of real-world interactions. The established image-based world models and policies have shown prior success, but lack robust geometric information that requires consistent spatial and physical understanding of the three-dimensional world, even pre-trained on internet-scale video sources. To this end, we propose a novel branch of world model named Gaussian World Model (GWM) for robotic manipulation, which reconstructs the future state by inferring the propagation of Gaussian primitives under the effect of robot actions. At its core is a latent Diffusion Transformer (DiT) combined with a 3D variational autoencoder, enabling fine-grained scene-level future state reconstruction with Gaussian Splatting. GWM can not only enhance the visual representation for imitation learning agent by self-supervised future prediction training, but can serve as a neural simulator that supports model-based reinforcement learning. Both simulated and real-world experiments depict that GWM can precisely predict future scenes conditioned on diverse robot actions, and can be further utilized to train policies that outperform the state-of-the-art by impressive margins, showcasing the initial data scaling potential of 3D world model.

3D世界模型机器人操作高斯点云扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。