用合成数据训练出媲美真实数据的通用机器人策略模型。
InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist Policy
- 构建全自动合成数据流水线,生成超63万条轨迹。
- 仅用合成数据训练的模型在49个仿真任务上达官方最优水平。
- 零样本迁移至真实场景,适合大规模机器人数据研究者。
近期研究探讨真实数据与合成数据对视觉-语言-动作(VLA)模型泛化能力的影响。尽管现有VLA模型在大规模真实机器人预训练中表现优异,但以往合成数据在规模上未能展现同等能力。本文首次证明,仅使用合成数据即可在预训练VLA模型时达到最强$π$-数据集的性能,揭示大规模仿真数据的巨大价值。所训练模型在多个挑战性任务中展现出意外的零样本模拟到现实迁移能力。本研究构建的合成数据集InternData-A1包含超过630,000条轨迹,总计7,433小时,覆盖4种机器人形态、18项技能、70个任务及227个场景,涵盖刚体、铰接体、可变形体和流体操作。数据通过高度自主、完全解耦且可组合的仿真流水线生成,支持长时序技能组合、灵活任务装配和异构机器人形态,几乎无需人工调参。采用与$π_0$相同架构,在InternData-A1上纯合成数据预训练的模型,在49个仿真任务、5个真实任务和4个长时序灵巧任务上表现与官方$π_0$相当。研究团队将开源该数据集及生成流水线,推动大规模机器人数据获取,降低具身智能研究的数据构建门槛。
原文摘要 · Abstract (English)
Recent works explore how real and synthetic data contribute to Vision-Language-Action (VLA) models' generalization. While current VLA models have shown the strong effectiveness of large-scale real-robot pre-training, synthetic data has not previously demonstrated comparable capability at scale. This paper provides the first evidence that synthetic data alone can match the performance of the strongest $π$-dataset in pre-training a VLA model, revealing the substantial value of large-scale simulation. The resulting model also exhibits surprisingly zero-shot sim-to-real transfer on several challenging tasks. Our synthetic dataset, InternData-A1, contains over 630k trajectories and 7,433 hours across 4 embodiments, 18 skills, 70 tasks, and 227 scenes, covering rigid, articulated, deformable, and fluid-object manipulation. It is generated through a highly autonomous, fully decoupled, and compositional simulation pipeline that enables long-horizon skill composition, flexible task assembly, and heterogeneous embodiments with minimal manual tuning. Using the same architecture as $π_0$, we pre-train a model entirely on InternData-A1 and find that it matches the official $π_0$ across 49 simulation tasks, 5 real-world tasks, and 4 long-horizon dexterous tasks. We release the dataset and will open-source the generation pipeline to broaden access to large-scale robotic data and to lower the barrier to scalable data creation for embodied AI research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。