arXiv:2503.18738cs.RO2025-03被引 52

无需标定即可生成带语义的机器人场景,提升模仿学习泛化能力。

RoboEngine: Plug-and-Play Robot Data Augmentation with Semantic Robot Segmentation and Background Generation

  • 基于语义分割与背景生成,实现即插即用的视觉数据增强。
  • 仅用单场景演示,在6个新场景中性能提升超200%。
  • 适合机器人仿真与强化学习研究者快速构建多样化训练环境。

视觉数据增强已成为提升模仿学习视觉鲁棒性的关键手段。然而,现有方法常受限于相机标定或受控环境(如绿幕)等前提条件。本文提出RoboEngine,首个即插即用的机器人视觉数据增强工具包。用户仅需几行代码,即可自动生成符合物理规律和任务需求的机器人场景。为此,我们构建了首个机器人场景分割数据集,开发了通用性强的高质量机器人分割模型,并微调了背景生成模型,三者共同构成开箱即用的核心组件。实验表明,仅使用单场景演示数据,RoboEngine即可在六个全新场景中实现任务泛化,性能较无增强基线提升超过200%。所有数据集、模型权重及工具包已公开:https://roboengine.github.io/

原文摘要 · Abstract (English)

Visual augmentation has become a crucial technique for enhancing the visual robustness of imitation learning. However, existing methods are often limited by prerequisites such as camera calibration or the need for controlled environments (e.g., green screen setups). In this work, we introduce RoboEngine, the first plug-and-play visual robot data augmentation toolkit. For the first time, users can effortlessly generate physics- and task-aware robot scenes with just a few lines of code. To achieve this, we present a novel robot scene segmentation dataset, a generalizable high-quality robot segmentation model, and a fine-tuned background generation model, which together form the core components of the out-of-the-box toolkit. Using RoboEngine, we demonstrate the ability to generalize robot manipulation tasks across six entirely new scenes, based solely on demonstrations collected from a single scene, achieving a more than 200% performance improvement compared to the no-augmentation baseline. All datasets, model weights, and the toolkit are released https://roboengine.github.io/

机器人数据增强模仿学习视觉生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。