用生成式AI合成农业图像,提升杂草识别模型训练效率
Generative AI-based Pipeline Architecture for Increasing Training Efficiency in Intelligent Weed Control Systems
- 结合SAM与Stable Diffusion生成真实场景的合成图像
- 仅用10%合成数据+90%真实数据,检测精度超纯真实数据训练
- 适合农业智能系统开发者,降低数据采集成本
在自动化作物保护任务如杂草控制、病害诊断和虫害监测中,深度学习展现出巨大潜力。然而,这些先进模型严重依赖高质量、多样化的数据集,而农业场景下的数据往往有限且获取成本高。传统数据增强虽能增加数据量,但难以涵盖真实世界中的多样性。本研究提出一种基于生成式AI的图像生成流水线,整合了零样本域适应的Segment Anything Model(SAM)与文本到图像的Stable Diffusion模型,生成能反映多样化真实条件的合成图像。我们使用轻量级YOLO模型评估这些合成数据集,通过不同比例的真实与合成数据组合,以mAP50和mAP50-95作为数据效率指标进行测试。结果显示,使用10%合成数据与90%真实数据训练的YOLO模型,在mAP50和mAP50-95上普遍优于仅使用真实数据训练的模型。该方法不仅减少对大规模真实数据的依赖,还提升了预测性能,为智能系统的感知模块实现持续自我优化提供了可能。
原文摘要 · Abstract (English)
In automated crop protection tasks such as weed control, disease diagnosis, and pest monitoring, deep learning has demonstrated significant potential. However, these advanced models rely heavily on high-quality, diverse datasets, often limited and costly in agricultural settings. Traditional data augmentation can increase dataset volume but usually lacks the real-world variability needed for robust training. This study presents a new approach for generating synthetic images to improve deep learning-based object detection models for intelligent weed control. Our GenAI-based image generation pipeline integrates the Segment Anything Model (SAM) for zero-shot domain adaptation with a text-to-image Stable Diffusion Model, enabling the creation of synthetic images that capture diverse real-world conditions. We evaluate these synthetic datasets using lightweight YOLO models, measuring data efficiency with mAP50 and mAP50-95 scores across varying proportions of real and synthetic data. Notably, YOLO models trained on datasets with 10% synthetic and 90% real images generally demonstrate superior mAP50 and mAP50-95 scores compared to those trained solely on real images. This approach not only reduces dependence on extensive real-world datasets but also enhances predictive performance. The integration of this approach opens opportunities for achieving continual self-improvement of perception modules in intelligent technical systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。