arXiv:2501.01895cs.ROcs.CV2025-01NeurIPS被引 62

EnerVerse让机器人预演未来操作空间,实现高效真实世界动作生成。

EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation

  • 基于分块自回归视频扩散模型,结合稀疏上下文记忆预测未来动作空间。
  • 在仿真与真实场景中均达当前最佳性能,8步动作生成仅需约280毫秒。
  • 适合机器人操控、具身智能研究者,尤其关注高效动作规划与跨域泛化。

我们提出EnerVerse,一种生成式机器人基础模型,用于构建和理解具身空间。EnerVerse采用分块自回归视频扩散框架,从指令预测未来具身空间,并通过稀疏上下文记忆支持长期推理。为建模三维机器人世界,采用多视角视频表示,提供丰富视角以应对运动模糊与三维定位挑战。此外,EnerVerse-D数据引擎结合生成建模与4D高斯点云渲染,形成自增强数据闭环,降低仿真到现实的差距。利用这些创新,EnerVerse通过策略头(EnerVerse-A)将4D世界表征转化为物理动作,在仿真与真实任务中均达到顶尖表现。为提升效率,EnerVerse-A复用首次去噪步骤特征并预测动作块,单张RTX 4090上每8步动作生成耗时约280毫秒。更多视频演示与数据样本详见项目页面。

原文摘要 · Abstract (English)

We introduce EnerVerse, a generative robotics foundation model that constructs and interprets embodied spaces. EnerVerse employs a chunk-wise autoregressive video diffusion framework to predict future embodied spaces from instructions, enhanced by a sparse context memory for long-term reasoning. To model the 3D robotics world, we adopt a multi-view video representation, providing rich perspectives to address challenges like motion ambiguity and 3D grounding. Additionally, EnerVerse-D, a data engine pipeline combining generative modeling with 4D Gaussian Splatting, forms a self-reinforcing data loop to reduce the sim-to-real gap. Leveraging these innovations, EnerVerse translates 4D world representations into physical actions via a policy head (EnerVerse-A), achieving state-of-the-art performance in both simulation and real-world tasks. For efficiency, EnerVerse-A reuses features from the first denoising step and predicts action chunks, achieving about 280 ms per 8-step action chunk on a single RTX 4090. Further video demos, dataset samples could be found in our project page.

机器人操控具身智能视频生成动作规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。