arXiv:2512.19632cs.CV2025-12被引 1

用生成模型解决农业图像数据少难题,提升作物识别与检测效果。

Generative diffusion models for agricultural AI: plant image generation, indoor-to-outdoor translation, and expert preference alignment

  • 用扩散模型生成逼真作物图像,支持文本控制
  • 合成图像使下游表型分类准确率提升,增强训练数据
  • 结合专家偏好微调,输出更符合专业需求

农业人工智能的成功依赖于大规模、多样且高质量的植物图像数据集,但真实田间环境下采集此类数据成本高、耗时长且受季节限制。本文探索基于扩散模型的生成方法,通过植物图像合成、室内到室外图像转换以及专家偏好对齐微调来应对这些挑战。首先,在带有描述的室内与室外植物图像上微调Stable Diffusion模型,生成油菜和大豆的文本条件真实图像;评估显示,合成图像在Inception Score、Frechet Inception Distance及下游表型分类任务中表现优异,有效扩充训练数据并提高精度。其次,采用DreamBooth文本反演与图像引导扩散技术,弥合高分辨率室内数据集与有限户外图像之间的差距,生成的转换图像显著提升了YOLOv8在杂草检测与分类中的性能。最后,构建基于专家评分的偏好微调框架,训练奖励模型并应用奖励加权更新,生成更稳定、更符合专家偏好的输出。三者结合展示了一条高效生成式农业AI数据管道的可行路径。

原文摘要 · Abstract (English)

The success of agricultural artificial intelligence depends heavily on large, diverse, and high-quality plant image datasets, yet collecting such data in real field conditions is costly, labor intensive, and seasonally constrained. This paper investigates diffusion-based generative modeling to address these challenges through plant image synthesis, indoor-to-outdoor translation, and expert preference aligned fine tuning. First, a Stable Diffusion model is fine tuned on captioned indoor and outdoor plant imagery to generate realistic, text conditioned images of canola and soybean. Evaluation using Inception Score, Frechet Inception Distance, and downstream phenotype classification shows that synthetic images effectively augment training data and improve accuracy. Second, we bridge the gap between high resolution indoor datasets and limited outdoor imagery using DreamBooth-based text inversion and image guided diffusion, generating translated images that enhance weed detection and classification with YOLOv8. Finally, a preference guided fine tuning framework trains a reward model on expert scores and applies reward weighted updates to produce more stable and expert aligned outputs. Together, these components demonstrate a practical pathway toward data efficient generative pipelines for agricultural AI.

生成模型农业AI图像合成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。