arXiv:2410.21153cs.CVcs.RO2024-10被引 8

用270万张合成图像训练出实时运行的机器人视觉检测模型

Synthetica: Large Scale Synthetic Data for Robot Perception

  • 基于光追渲染生成大规模合成数据,提升模型鲁棒性
  • 实现50-100Hz实时推理,速度比前代快9倍
  • 适用于工业定制物体,无需真实数据集

基于视觉的目标检测是机器人应用的关键基础,能提供环境中物体位置信息。这类模型需在不同光照、遮挡和视觉伪影条件下保持高可靠性,同时实现实时运行。收集并标注真实世界数据成本高昂,尤其针对定制物体(如工业物品),难以推广到真实场景。为此,我们提出Synthetica,一种用于训练鲁棒状态估计算法的大规模合成数据生成方法。本文聚焦目标检测任务,该任务可作为多数状态估计问题的前端。利用基于光真实的光线追踪渲染器生成数据,我们扩展了数据生成规模,共生成270万张图像,用于训练高精度实时检测变换模型。提出一系列渲染随机化与训练时数据增强技术,有助于提升视觉任务的模拟到现实迁移性能。在目标检测任务上达到当前最优表现,且检测器运行速度达50-100Hz,比先前最先进方法快9倍。进一步通过定制物体的真实场景应用管道,验证了该训练方法对机器人应用的有效性。本工作强调了大规模合成数据生成对稳健模拟到现实迁移的重要性,并实现了最快实时推理速度。视频与补充资料见:https://sites.google.com/view/synthetica-vision。

原文摘要 · Abstract (English)

Vision-based object detectors are a crucial basis for robotics applications as they provide valuable information about object localisation in the environment. These need to ensure high reliability in different lighting conditions, occlusions, and visual artifacts, all while running in real-time. Collecting and annotating real-world data for these networks is prohibitively time consuming and costly, especially for custom assets, such as industrial objects, making it untenable for generalization to in-the-wild scenarios. To this end, we present Synthetica, a method for large-scale synthetic data generation for training robust state estimators. This paper focuses on the task of object detection, an important problem which can serve as the front-end for most state estimation problems, such as pose estimation. Leveraging data from a photorealistic ray-tracing renderer, we scale up data generation, generating 2.7 million images, to train highly accurate real-time detection transformers. We present a collection of rendering randomization and training-time data augmentation techniques conducive to robust sim-to-real performance for vision tasks. We demonstrate state-of-the-art performance on the task of object detection while having detectors that run at 50-100Hz which is 9 times faster than the prior SOTA. We further demonstrate the usefulness of our training methodology for robotics applications by showcasing a pipeline for use in the real world with custom objects for which there do not exist prior datasets. Our work highlights the importance of scaling synthetic data generation for robust sim-to-real transfer while achieving the fastest real-time inference speeds. Videos and supplementary information can be found at this URL: https://sites.google.com/view/synthetica-vision.

合成数据机器人感知实时检测光追渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。