用大模型生成高保真机器人场景,支持大规模训练与自动评估
Genie Sim 3.0 : A High-Fidelity Comprehensive Simulation Platform for Humanoid Robot
- 用大语言模型根据自然语言指令自动生成复杂仿真环境
- 构建首个基于大模型的自动化评估基准,支持超200项任务测试
- 开源超1万小时合成数据,实现零样本跨仿真到现实迁移
机器人学习模型的鲁棒性与泛化能力高度依赖大规模、多样化的训练数据和可靠评估基准。真实世界数据采集成本高昂且难以扩展,现有仿真基准普遍存在碎片化、范围狭窄或保真度不足的问题,难以实现有效的仿真到现实迁移。为此,我们提出 Genie Sim 3.0——一个统一的人形机器人操作仿真平台。引入 Genie Sim Generator,一种基于大语言模型(LLM)的工具,可从自然语言指令生成高保真场景。其核心优势在于快速、多维度的泛化能力,支持多样化环境生成,推动可扩展的数据收集与稳健策略评估。我们提出首个利用大模型实现自动化评估的基准,通过 LLM 大规模生成评估场景,并借助视觉-语言模型(VLM)建立自动化评估流程。同时,我们发布一个开源数据集,包含超过 10,000 小时的合成数据,覆盖 200 多项任务。系统实验验证了该数据集在受控条件下具备零样本仿真到现实迁移能力,表明合成数据可有效替代真实数据用于规模化策略训练。
原文摘要 · Abstract (English)
The development of robust and generalizable robot learning models is critically contingent upon the availability of large-scale, diverse training data and reliable evaluation benchmarks. Collecting data in the physical world poses prohibitive costs and scalability challenges, and prevailing simulation benchmarks frequently suffer from fragmentation, narrow scope, or insufficient fidelity to enable effective sim-to-real transfer. To address these challenges, we introduce Genie Sim 3.0, a unified simulation platform for robotic manipulation. We present Genie Sim Generator, a large language model (LLM)-powered tool that constructs high-fidelity scenes from natural language instructions. Its principal strength resides in rapid and multi-dimensional generalization, facilitating the synthesis of diverse environments to support scalable data collection and robust policy evaluation. We introduce the first benchmark that pioneers the application of LLM for automated evaluation. It leverages LLM to mass-generate evaluation scenarios and employs Vision-Language Model (VLM) to establish an automated assessment pipeline. We also release an open-source dataset comprising more than 10,000 hours of synthetic data across over 200 tasks. Through systematic experimentation, we validate the robust zero-shot sim-to-real transfer capability of our open-source dataset, demonstrating that synthetic data can server as an effective substitute for real-world data under controlled conditions for scalable policy training. For code and dataset details, please refer to: https://github.com/AgibotTech/genie_sim.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。