arXiv:2511.19861cs.CVcs.RO2025-11被引 47

用世界模型生成逼真交互数据,让机器人无需真实训练就能学会任务。

GigaWorld-0: World Models as Data Engine to Empower Embodied AI

  • 构建双模块框架,融合视频与3D生成,实现可控、物理真实的虚拟交互数据
  • 在GigaBrain-0上训练的模型零实机交互即达成强泛化性能,任务成功率显著提升
  • 适合研究具身智能、数据生成与机器人学习的开发者快速获取高质量仿真数据

世界模型正成为可扩展、数据高效的具身智能基础范式。本文提出GigaWorld-0,一个专为视觉-语言-动作(VLA)学习设计的统一世界模型框架,作为数据引擎。该框架包含两个协同组件:GigaWorld-0-Video利用大规模视频生成,在精细控制外观、相机视角和动作语义下产出多样、纹理丰富、时序连贯的具身序列;GigaWorld-0-3D结合3D生成建模、3D高斯泼溅重建、可微分物理系统辨识与可执行运动规划,确保几何一致性与物理真实性。两者联合优化,实现了视觉吸引人、空间一致、物理合理且指令对齐的大规模具身交互数据合成。通过高效GigaTrain框架(采用FP8精度与稀疏注意力),大幅降低内存与计算需求,使大规模训练成为可能。全面评估表明,GigaWorld-0在多维度生成高质量、多样且可控的数据。关键的是,基于其生成数据训练的VLA模型(如GigaBrain-0)在真实机器人上表现强劲,无需任何真实交互即可显著提升泛化能力与任务成功率。

原文摘要 · Abstract (English)

World models are emerging as a foundational paradigm for scalable, data-efficient embodied AI. In this work, we present GigaWorld-0, a unified world model framework designed explicitly as a data engine for Vision-Language-Action (VLA) learning. GigaWorld-0 integrates two synergistic components: GigaWorld-0-Video, which leverages large-scale video generation to produce diverse, texture-rich, and temporally coherent embodied sequences under fine-grained control of appearance, camera viewpoint, and action semantics; and GigaWorld-0-3D, which combines 3D generative modeling, 3D Gaussian Splatting reconstruction, physically differentiable system identification, and executable motion planning to ensure geometric consistency and physical realism. Their joint optimization enables the scalable synthesis of embodied interaction data that is visually compelling, spatially coherent, physically plausible, and instruction-aligned. Training at scale is made feasible through our efficient GigaTrain framework, which exploits FP8-precision and sparse attention to drastically reduce memory and compute requirements. We conduct comprehensive evaluations showing that GigaWorld-0 generates high-quality, diverse, and controllable data across multiple dimensions. Critically, VLA model (e.g., GigaBrain-0) trained on GigaWorld-0-generated data achieve strong real-world performance, significantly improving generalization and task success on physical robots without any real-world interaction during training.

具身智能世界模型数据生成机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。