EnerVerse让机器人预演未来操作空间,实现高效真实世界动作生成。
EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation
- 基于分块自回归视频扩散模型,结合稀疏上下文记忆预测未来动作空间。
- 在仿真与真实场景中均达当前最佳性能,8步动作生成仅需约280毫秒。
- 适合机器人操控、具身智能研究者,尤其关注高效动作规划与跨域泛化。
我们提出EnerVerse,一种生成式机器人基础模型,用于构建和理解具身空间。EnerVerse采用分块自回归视频扩散框架,从指令预测未来具身空间,并通过稀疏上下文记忆支持长期推理。为建模三维机器人世界,采用多视角视频表示,提供丰富视角以应对运动模糊与三维定位挑战。此外,EnerVerse-D数据引擎结合生成建模与4D高斯点云渲染,形成自增强数据闭环,降低仿真到现实的差距。利用这些创新,EnerVerse通过策略头(EnerVerse-A)将4D世界表征转化为物理动作,在仿真与真实任务中均达到顶尖表现。为提升效率,EnerVerse-A复用首次去噪步骤特征并预测动作块,单张RTX 4090上每8步动作生成耗时约280毫秒。更多视频演示与数据样本详见项目页面。
原文摘要 · Abstract (English)
We introduce EnerVerse, a generative robotics foundation model that constructs and interprets embodied spaces. EnerVerse employs a chunk-wise autoregressive video diffusion framework to predict future embodied spaces from instructions, enhanced by a sparse context memory for long-term reasoning. To model the 3D robotics world, we adopt a multi-view video representation, providing rich perspectives to address challenges like motion ambiguity and 3D grounding. Additionally, EnerVerse-D, a data engine pipeline combining generative modeling with 4D Gaussian Splatting, forms a self-reinforcing data loop to reduce the sim-to-real gap. Leveraging these innovations, EnerVerse translates 4D world representations into physical actions via a policy head (EnerVerse-A), achieving state-of-the-art performance in both simulation and real-world tasks. For efficiency, EnerVerse-A reuses features from the first denoising step and predicts action chunks, achieving about 280 ms per 8-step action chunk on a single RTX 4090. Further video demos, dataset samples could be found in our project page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。