用百亿帧仿真数据训练抓取大模型,实现零样本泛化。
GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data
- 用仿真生成百亿帧数据,结合视觉语言动作联合训练。
- 在真实和仿真任务中实现零样本抓取成功率超80%。
- 适合机器人抓取、具身智能研究者快速部署使用。
具身基础模型因其零样本泛化、可扩展性及少量微调即可适应新任务而受到关注。然而现有模型严重依赖真实世界数据,采集成本高且耗时。合成数据提供了一种低成本替代方案,但其潜力尚未充分挖掘。为此,我们探索了完全基于大规模合成动作数据训练视觉-语言-动作(VLA)模型的可行性。我们构建了SynGrasp-1B,一个包含百亿帧的机器人抓取数据集,通过逼真渲染与广泛域随机化生成。基于此,我们提出GraspVLA,一种在大规模合成动作数据上预训练的抓取基础模型。GraspVLA将自回归感知任务与基于流匹配的动作生成整合进统一的思维链流程中,支持在合成动作数据与互联网语义数据上联合训练。该设计有助于缓解仿真到现实的差距,并促进所学动作向更广泛互联网覆盖物体的迁移,实现开放词汇抓取泛化。在真实世界与仿真基准上的大量评估表明,GraspVLA具备先进的零样本泛化能力与少量样本适应特定人类偏好能力。我们将公开SynGrasp-1B数据集与预训练权重,以推动社区发展。
原文摘要 · Abstract (English)
Embodied foundation models are gaining increasing attention for their zero-shot generalization, scalability, and adaptability to new tasks through few-shot post-training. However, existing models rely heavily on real-world data, which is costly and labor-intensive to collect. Synthetic data offers a cost-effective alternative, yet its potential remains largely underexplored. To bridge this gap, we explore the feasibility of training Vision-Language-Action models entirely with large-scale synthetic action data. We curate SynGrasp-1B, a billion-frame robotic grasping dataset generated in simulation with photorealistic rendering and extensive domain randomization. Building on this, we present GraspVLA, a VLA model pretrained on large-scale synthetic action data as a foundational model for grasping tasks. GraspVLA integrates autoregressive perception tasks and flow-matching-based action generation into a unified Chain-of-Thought process, enabling joint training on synthetic action data and Internet semantics data. This design helps mitigate sim-to-real gaps and facilitates the transfer of learned actions to a broader range of Internet-covered objects, achieving open-vocabulary generalization in grasping. Extensive evaluations across real-world and simulation benchmarks demonstrate GraspVLA's advanced zero-shot generalizability and few-shot adaptability to specific human preferences. We will release SynGrasp-1B dataset and pre-trained weights to benefit the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。