arXiv:2510.19430cs.ROcs.CV2025-10被引 42

用世界模型生成数据训练机器人模型,减少真实数据依赖。

GigaBrain-0: A World Model-Powered Vision-Language-Action Model

  • 用世界模型生成视频、视角转移等数据替代真实采集
  • 在复杂操作任务中实现跨场景强泛化,性能显著提升
  • 轻量版可部署于Jetson设备,适合边缘智能应用

通用机器人视觉-语言-动作(VLA)模型训练通常依赖大规模真实机器人数据,收集成本高、效率低,严重制约系统可扩展性和泛化能力。为此,我们提出GigaBrain-0,一种由世界模型生成数据(如视频生成、真实间迁移、人体迁移、视角迁移、仿真到现实迁移)驱动的新型VLA基础模型。通过大规模生成多样化数据,显著降低对真实机器人数据的依赖,同时提升跨任务泛化能力。方法进一步通过RGBD输入建模与具身链式思维(CoT)监督,增强模型对空间几何、物体状态及长程依赖的推理能力,显著提升实际操作中的鲁棒性。大量实验表明,GigaBrain-0在灵巧操作、长时序任务和移动操作任务中均表现出色,对外观(如纹理、颜色)、物体位置和相机视角变化具有强鲁棒性。此外,我们还推出GigaBrain-0-Small,一个专为NVIDIA Jetson AGX Orin等设备优化的轻量版本。

原文摘要 · Abstract (English)

Training Vision-Language-Action (VLA) models for generalist robots typically requires large-scale real-world robot data, which is expensive and time-consuming to collect. The inefficiency of physical data collection severely limits the scalability, and generalization capacity of current VLA systems. To address this challenge, we introduce GigaBrain-0, a novel VLA foundation model empowered by world model-generated data (e.g., video generation, real2real transfer, human transfer, view transfer, sim2real transfer data). By leveraging world models to generate diverse data at scale, GigaBrain-0 significantly reduces reliance on real robot data while improving cross-task generalization. Our approach further improves policy robustness through RGBD input modeling and embodied Chain-of-Thought (CoT) supervision, enabling the model to reason about spatial geometry, object states, and long-horizon dependencies during task execution. This leads to substantial gains in real-world performance on dexterous, long-horizon, and mobile manipulation tasks. Extensive experiments demonstrate that GigaBrain-0 achieves superior generalization across variations in appearances (e.g., textures, colors), object placements, and camera viewpoints. Additionally, we present GigaBrain-0-Small, an optimized lightweight variant designed to run efficiently on devices such as the NVIDIA Jetson AGX Orin.

机器人世界模型VLA轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。