自动生成可执行的3D仿真环境,支持智能体训练与真实机器人部署。
EmbodiedGen V2: An Agentic, Simulation-Ready 3D World Engine for Embodied AI

- 统一表示跨模拟器资产与交互功能,实现可编辑的生成式仿真流水线。
- 生成环境98.6%碰撞成功率,83.3%任务场景无需修改即可直接使用。
- 适用于机器人操控、导航及强化学习,支持从仿真到现实的迁移。
我们提出EmbodiedGen V2,一个面向具身智能的生成式3D世界引擎,用于构建可执行的策略就绪环境。尽管模拟器就绪的3D资产生成已快速进展,但将这些资产组装成策略就绪的任务环境仍主要依赖人工,限制了闭环学习的可扩展性。EmbodiedGen V2通过统一的模拟器就绪表示,连接跨模拟器资产、交互属性、任务驱动世界、大规模多房间场景以及状态化Vibe编码,构建了一个可生成、可编辑、可复用的仿真流水线。生成环境支持操作、导航、移动操作、跨模拟器部署及具身策略训练。评估显示,资产流水线的人类接受率达96.5%,碰撞成功率达98.6%,83.3%的任务驱动世界可直接用于下游仿真而无需手动修改。基于生成环境的在线强化学习使仿真成功率从9.7%提升至79.8%,并迁移到真实机器人后任务成功率从21.7%提升至75.0%。这些结果确立了EmbodiedGen V2作为具身策略训练、评估与部署的可扩展仿真基础设施。
原文摘要 · Abstract (English)
We present EmbodiedGen V2, a generative 3D world engine for building executable policy-ready environments for embodied intelligence. Sim-ready 3D asset generation has advanced rapidly, yet assembling such assets into policy-ready task environments remains largely manual, limiting scalable closed-loop learning. EmbodiedGen V2 addresses this gap through a unified sim-ready representation that connects cross-simulator assets, interaction affordances, task-driven worlds, large-scale multi-room scenes, and stateful Vibe Coding into a generative, editable, and reusable simulation pipeline. The generated environments support manipulation, navigation, mobile manipulation, cross-simulator deployment, and embodied policy training. In evaluation, the asset pipeline achieves 96.5% human acceptance and 98.6% collision success, and 83.3% of task-driven worlds are directly usable for downstream simulation without manual modification. Online reinforcement learning with generated environments further improves simulation success from 9.7% to 79.8%, and transfers to real robots with task success increasing from 21.7% to 75.0%. These results establish EmbodiedGen V2 as scalable simulation infrastructure for training, evaluating, and deploying embodied policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。