arXiv:2506.07539cs.CVcs.AI2025-06ICRA被引 10

用合成数据提升工业检测精度,仅靠仿真训练就达顶尖效果

Domain Randomization for Object Detection in Manufacturing Applications using Synthetic Data: A Comprehensive Study

  • 系统化随机化物体、光照、背景等要素生成逼真合成数据
  • 在真实工业数据集上实现94.1%~99.5%的检测准确率
  • 适合做工业视觉检测的算法研发者参考

本文研究了在制造场景中生成合成数据时域随机化的关键因素。提出一个涵盖物体属性、背景、光照、相机参数及后处理的综合数据生成流程,并构建了包含15类工业零件的SIP15-OD数据集,用于多场景测试。同时使用公开的机器人应用工业数据集进行验证。实验揭示材料特性、渲染方式、后处理和干扰物是影响性能的关键因素。基于这些因素优化的合成数据,使YOLOv8模型仅在合成数据上训练即达到优异表现:在机器人数据集上mAP@50达96.4%,在SIP15-OD三个应用场景中分别达到94.1%、99.5%、95.3%。结果表明该方法能有效逼近真实数据分布。

原文摘要 · Abstract (English)

This paper addresses key aspects of domain randomization in generating synthetic data for manufacturing object detection applications. To this end, we present a comprehensive data generation pipeline that reflects different factors: object characteristics, background, illumination, camera settings, and post-processing. We also introduce the Synthetic Industrial Parts Object Detection dataset (SIP15-OD) consisting of 15 objects from three industrial use cases under varying environments as a test bed for the study, while also employing an industrial dataset publicly available for robotic applications. In our experiments, we present more abundant results and insights into the feasibility as well as challenges of sim-to-real object detection. In particular, we identified material properties, rendering methods, post-processing, and distractors as important factors. Our method, leveraging these, achieves top performance on the public dataset with Yolov8 models trained exclusively on synthetic data; mAP@50 scores of 96.4% for the robotics dataset, and 94.1%, 99.5%, and 95.3% across three of the SIP15-OD use cases, respectively. The results showcase the effectiveness of the proposed domain randomization, potentially covering the distribution close to real data for the applications.

工业检测合成数据域随机化目标检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。