用图文生成模型增强机器人数据,提升真实场景泛化能力。
Semantically Controllable Augmentations for Generalizable Robot Learning
- 用预训练图文模型生成语义可控的合成数据
- 在模拟和真实厨房场景中实现高效泛化
- 无需额外人工成本,适合大规模机器人学习
机器人在未见的真实场景中实现操作泛化,需要训练时接触多样化数据。然而,由于操作成本高昂,收集大规模真实数据集不切实际。为克服这一挑战,本文提出利用图像-文本生成模型作为数据源。这些模型基于网络抓取的大规模数据预训练,涵盖远超机器人直接经验的真实世界场景,可合成新颖的合成体验,让机器人代理在无额外成本下接触更多世界先验知识。本文提出一种生成式数据增强框架,实现语义可控的数据扩充,快速扩展机器人数据集,并引入丰富变化,支持真实世界泛化。基于多样化的数据增强,我们展示了可在仿真和未见真实环境(如厨房、桌面)中训练并部署可扩展的机器人操作策略。实验表明,图像-文本生成模型在多种真实机器人应用中有效,该框架为机器人学习提供了低成本、可扩展的泛化路径。
原文摘要 · Abstract (English)
Generalization to unseen real-world scenarios for robot manipulation requires exposure to diverse datasets during training. However, collecting large real-world datasets is intractable due to high operational costs. For robot learning to generalize despite these challenges, it is essential to leverage sources of data or priors beyond the robot's direct experience. In this work, we posit that image-text generative models, which are pre-trained on large corpora of web-scraped data, can serve as such a data source. These generative models encompass a broad range of real-world scenarios beyond a robot's direct experience and can synthesize novel synthetic experiences that expose robotic agents to additional world priors aiding real-world generalization at no extra cost. In particular, our approach leverages pre-trained generative models as an effective tool for data augmentation. We propose a generative augmentation framework for semantically controllable augmentations and rapidly multiplying robot datasets while inducing rich variations that enable real-world generalization. Based on diverse augmentations of robot data, we show how scalable robot manipulation policies can be trained and deployed both in simulation and in unseen real-world environments such as kitchens and table-tops. By demonstrating the effectiveness of image-text generative models in diverse real-world robotic applications, our generative augmentation framework provides a scalable and efficient path for boosting generalization in robot learning at no extra human cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。