用3D生成合成火焰图训练检测模型,效果优于纯真实或纯合成数据。
Synthetic imagery for fuzzy object detection: A comparative study
- 基于3D模型生成合成火焰图像并自动标注
- 混合训练集使模型在多场景火焰检测上表现最佳
- 大幅降低标注成本,适合模糊物体检测研究者
模糊物体检测是计算机视觉中的挑战性课题。火焰、烟雾、雾气和蒸汽等模糊物体在视觉特征上更具复杂性,表现为边缘模糊、形状可变、透明度不一及体积变化。构建平衡且多样化的数据集并进行精确标注对提升机器学习模型性能至关重要,但该过程仍高度依赖人工。本文提出一种基于3D模型生成并自动标注全合成火焰图像的方法,用于训练目标检测模型。对比分析了仅使用合成图像、仅使用真实图像以及混合图像训练的模型性能与效率。结果表明,合成数据在火焰检测中有效;当测试数据覆盖更广的真实火焰场景时,模型性能进一步提升。混合训练集下的模型优于仅使用真实或仅使用合成数据训练的模型,能更好检测多种火焰。该自动化合成与标注方法显著降低了创建针对模糊物体检测的视觉模型的时间与成本。
原文摘要 · Abstract (English)
The fuzzy object detection is a challenging field of research in computer vision (CV). Distinguishing between fuzzy and non-fuzzy object detection in CV is important. Fuzzy objects such as fire, smoke, mist, and steam present significantly greater complexities in terms of visual features, blurred edges, varying shapes, opacity, and volume compared to non-fuzzy objects such as trees and cars. Collection of a balanced and diverse dataset and accurate annotation is crucial to achieve better ML models for fuzzy objects, however, the task of collection and annotation is still highly manual. In this research, we propose and leverage an alternative method of generating and automatically annotating fully synthetic fire images based on 3D models for training an object detection model. Moreover, the performance, and efficiency of the trained ML models on synthetic images is compared with ML models trained on real imagery and mixed imagery. Findings proved the effectiveness of the synthetic data for fire detection, while the performance improves as the test dataset covers a broader spectrum of real fires. Our findings illustrates that when synthetic imagery and real imagery is utilized in a mixed training set the resulting ML model outperforms models trained on real imagery as well as models trained on synthetic imagery for detection of a broad spectrum of fires. The proposed method for automating the annotation of synthetic fuzzy objects imagery carries substantial implications for reducing both time and cost in creating computer vision models specifically tailored for detecting fuzzy objects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。