arXiv:2605.20752cs.RO2026-05被引 2

用3D高斯模型提升机器人操作的精准度与效率

GaussianDream: A Feed-Forward 3D Gaussian World Model for Robotic Manipulation

论文配图:GaussianDream: A Feed-Forward 3D Gaussian World Model for Robotic Manipulation
图 1 · 摘自论文原文
  • 引入可学习的高斯查询,实时捕捉3D空间结构和短期未来变化
  • 在LIBERO上达98.4%成功率,实机任务也达50.0%
  • 推理时无需重建或预测,适合部署于实时控制场景

视觉-语言-动作(VLA)策略通过将预训练视觉-语言模型的语义先验迁移至动作生成,推动了语言控制下的机器人操作发展。然而,标准的动作模仿学习常缺乏对显式3D空间信息、密集几何监督及未来环境演化的建模,而这些对精确交互至关重要。为此,我们提出 extbf{GaussianDream},一种前馈式3D高斯世界模型插件。具体地,在编码器中引入可学习的GaussianDream查询,使模型能够捕获当前帧的3D空间结构和短时未来演化。训练时,潜在的GaussianDream前缀经由静态重建头和未来预测头处理,生成当前3D高斯场景状态与未来高斯演化状态。当前分支由RGB渲染和深度图监督,未来分支则使用未来RGB、深度及伪3D场景流信号。推理时,GaussianDream丢弃所有辅助头,仅保留学习到的前缀以条件化动作生成,无需测试时的高斯重建或未来预测。实验表明,GaussianDream在多个机器人操作基准上达到顶尖性能:LIBERO上达98.4%,RoboCasa Human-50为54.8%,真实机器人任务为50.0%。相比现有3D增强型VLA方法,GaussianDream在保持高精度的同时,推理效率高于基于视频的世界模型方法。

原文摘要 · Abstract (English)

Vision-language-action (VLA) policies have advanced language-conditioned robotic manipulation by transferring semantic priors from pretrained vision-language models to action generation. However, standard action-imitation learning often lacks sufficient modeling of explicit 3D spatial information, dense geometric supervision, and future environment evolution, all critical for precise robotic interaction. To address this, we propose \textbf{GaussianDream}, a feed-forward 3D Gaussian world-model plug-in. Specifically, we introduce learnable GaussianDream Queries in the encoder, enabling the model to capture current-frame 3D spatial structure and short-horizon future evolution. During training, the latent GaussianDream prefix is processed by a static reconstruction head and a future prediction head to produce current 3D Gaussian scene states and future Gaussian evolution states. The current branch is supervised by RGB rendering and depth, while the future branch uses future RGB, depth, and pseudo 3D scene-flow signals. During inference, GaussianDream discards all auxiliary heads and retains only the learned prefix to condition action generation, without test-time Gaussian reconstruction or future prediction. Experimental results demonstrate that GaussianDream achieves state-of-the-art performance across multiple robotic manipulation benchmarks, reaching \textbf{98.4\%} on LIBERO, \textbf{54.8\%} on RoboCasa Human-50, and \textbf{50.0\%} on real-robot tasks. Compared with existing 3D-enhanced VLA methods, GaussianDream achieves strong accuracy while providing higher inference efficiency than video-based world-model approaches.

机器人操作3D建模高斯表示高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。