用Midjourney生成1.2万张工地工人图像,解决数据少难题。
Synthesizing Reality: Leveraging the Generative AI-Powered Platform Midjourney for Construction Worker Detection
- 用3000个提示词生成1.2万张逼真合成图,提升数据多样性。
- 在真实数据集上检测平均精度达0.937(IoU=0.5)。
- 适合缺乏标注数据的工业视觉场景,尤其工地安全监控。
尽管深度神经网络(DNN)在视觉人工智能方面取得显著进展,但施工领域仍面临数据多样性与数量不足的挑战。本研究提出一种针对施工工人检测的新型图像合成方法,利用生成式AI平台Midjourney生成12,000张合成图像,通过3000个不同提示词实现高真实感与多样性。经人工标注后,这些图像构成DNN训练数据集。在真实施工图像数据集上的评估显示,模型在交并比(IoU)阈值为0.5和0.5至0.95时,平均精度(AP)分别达到0.937和0.642。值得注意的是,模型在合成数据集上表现接近完美,对应阈值下AP分别为0.994和0.919。结果表明,生成式AI在缓解DNN训练数据稀缺问题上具有潜力,但也暴露其局限性。
原文摘要 · Abstract (English)
While recent advancements in deep neural networks (DNNs) have substantially enhanced visual AI's capabilities, the challenge of inadequate data diversity and volume remains, particularly in construction domain. This study presents a novel image synthesis methodology tailored for construction worker detection, leveraging the generative-AI platform Midjourney. The approach entails generating a collection of 12,000 synthetic images by formulating 3000 different prompts, with an emphasis on image realism and diversity. These images, after manual labeling, serve as a dataset for DNN training. Evaluation on a real construction image dataset yielded promising results, with the model attaining average precisions (APs) of 0.937 and 0.642 at intersection-over-union (IoU) thresholds of 0.5 and 0.5 to 0.95, respectively. Notably, the model demonstrated near-perfect performance on the synthetic dataset, achieving APs of 0.994 and 0.919 at the two mentioned thresholds. These findings reveal both the potential and weakness of generative AI in addressing DNN training data scarcity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。