用多模态生成模型高效合成作物病害图像,兼顾质量与计算效率。
PhytoSynth: Leveraging Multi-modal Generative Models for Crop Disease Data Generation with Novel Benchmarking and Prompt Engineering Approach
- 基于文本到图像的生成模型,结合微调技术提升合成效果。
- 仅用36张真实样本1.5小时生成500张合成图像,能耗仅0.002kWh/张。
- 首次提供农业生成任务的算力基准,适合数据稀缺场景应用。
田间大规模采集作物病害图像成本高、耗时长。生成模型(GMs)可作为替代方案,生成接近真实图像的合成样本。然而,现有研究主要依赖生成对抗网络(GAN)进行图像到图像转换,且缺乏对农业场景中计算需求的全面分析。本文探索了一种多模态文本到图像的方法,用于生成合成作物病害图像,并首次在该领域提供计算性能基准。我们训练了三种Stable Diffusion变体——SDXL、SD3.5M(中等)和SD3.5L(大),并采用Dreambooth与低秩适应(LoRA)微调技术以增强泛化能力。实验表明,SD3.5M表现最优,推理时平均内存占用18 GB,功耗180 W,总能耗1.02 kWh生成500张图像(每张0.002 kWh)。结果证明,仅需36张实地样本,即可在1.5小时内生成500张高质量合成图像。建议在作物病害数据生成中优先使用SD3.5M。
原文摘要 · Abstract (English)
Collecting large-scale crop disease images in the field is labor-intensive and time-consuming. Generative models (GMs) offer an alternative by creating synthetic samples that resemble real-world images. However, existing research primarily relies on Generative Adversarial Networks (GANs)-based image-to-image translation and lack a comprehensive analysis of computational requirements in agriculture. Therefore, this research explores a multi-modal text-to-image approach for generating synthetic crop disease images and is the first to provide computational benchmarking in this context. We trained three Stable Diffusion (SD) variants-SDXL, SD3.5M (medium), and SD3.5L (large)-and fine-tuned them using Dreambooth and Low-Rank Adaptation (LoRA) fine-tuning techniques to enhance generalization. SD3.5M outperformed the others, with an average memory usage of 18 GB, power consumption of 180 W, and total energy use of 1.02 kWh/500 images (0.002 kWh per image) during inference task. Our results demonstrate SD3.5M's ability to generate 500 synthetic images from just 36 in-field samples in 1.5 hours. We recommend SD3.5M for efficient crop disease data generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。