arXiv:2508.05635cs.ROcs.CV2025-08被引 125

一个统一平台让机器人学会按指令操作,生成真实动作并自动评估。

Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation

  • 用视频扩散模型学习真实机器人的动作规律,建模空间时间语义关系。
  • 轻量解码器将抽象表示转为可执行动作,少标注也能跨机型通用。
  • 自带仿真与评测工具,适合做通用机器人智能研究的开发者使用。

我们提出Genie Envisioner(GE),一个面向机器人操作的统一世界基础平台,将策略学习、评估与仿真集成于单一视频生成框架中。核心GE-Base是一个大规模、指令条件化的视频扩散模型,以结构化潜在空间捕捉真实世界机器人交互中的空间、时间与语义动态。基于此,GE-Act通过轻量级流匹配解码器将潜在表示映射为可执行动作轨迹,实现低监督下多种机器人形态的精准且可泛化的策略推理。为支持可扩展的训练与评估,GE-Sim作为动作条件化的神经仿真器,生成高保真回放数据以支持闭环策略开发。平台还配备EWMBench基准套件,用于衡量视觉保真度、物理一致性及指令-动作对齐程度。上述组件共同构建了一个可扩展、实用的指令驱动型通用具身智能基础平台。所有代码、模型与基准将公开发布。

原文摘要 · Abstract (English)

We introduce Genie Envisioner (GE), a unified world foundation platform for robotic manipulation that integrates policy learning, evaluation, and simulation within a single video-generative framework. At its core, GE-Base is a large-scale, instruction-conditioned video diffusion model that captures the spatial, temporal, and semantic dynamics of real-world robotic interactions in a structured latent space. Built upon this foundation, GE-Act maps latent representations to executable action trajectories through a lightweight, flow-matching decoder, enabling precise and generalizable policy inference across diverse embodiments with minimal supervision. To support scalable evaluation and training, GE-Sim serves as an action-conditioned neural simulator, producing high-fidelity rollouts for closed-loop policy development. The platform is further equipped with EWMBench, a standardized benchmark suite measuring visual fidelity, physical consistency, and instruction-action alignment. Together, these components establish Genie Envisioner as a scalable and practical foundation for instruction-driven, general-purpose embodied intelligence. All code, models, and benchmarks will be released publicly.

机器人操作视频生成具身智能扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。