用文本引导扩散模型生成适配果园的标注图像,提升植株检测精度。
D4: Text-guided diffusion model-based domain adaptive data augmentation for vineyard shoot detection
- 基于文本引导扩散模型,从少量标注数据生成带目标域背景的新图像。
- 在边界框检测中平均精度提升28.65%,关键点检测提升13.73%。
- 适合农业视觉任务中数据少、场景多变时的模型泛化优化。
在农业领域,基于目标检测模型的植物表型分析日益受到关注。然而,由于标注难度大及域间多样性高,获取通用且高精度模型所需的训练数据极为困难。此外,跨作物间难以迁移训练数据,尽管已有针对特定环境、条件或作物的机器学习模型,但难以在实际田间广泛应用。本文提出一种面向葡萄园枝条检测的生成式数据增强方法(D4)。D4利用从无人地面车辆等手段采集的视频数据中提取的大量原始图像,结合少量标注数据,基于预训练的文本引导扩散模型生成新标注图像。该方法在保留目标检测所需标注信息的同时,使背景信息适配目标域。实验表明,该方法在边界框检测任务中将平均精度提升最高达28.65%,在关键点检测任务中平均精度提升最高达13.73%。D4有望同时解决农业训练数据生成的成本与域多样性问题,并提升检测模型的泛化性能。
原文摘要 · Abstract (English)
In an agricultural field, plant phenotyping using object detection models is gaining attention. However, collecting the training data necessary to create generic and high-precision models is extremely challenging due to the difficulty of annotation and the diversity of domains. Furthermore, it is difficult to transfer training data across different crops, and although machine learning models effective for specific environments, conditions, or crops have been developed, they cannot be widely applied in actual fields. In this study, we propose a generative data augmentation method (D4) for vineyard shoot detection. D4 uses a pre-trained text-guided diffusion model based on a large number of original images culled from video data collected by unmanned ground vehicles or other means, and a small number of annotated datasets. The proposed method generates new annotated images with background information adapted to the target domain while retaining annotation information necessary for object detection. In addition, D4 overcomes the lack of training data in agriculture, including the difficulty of annotation and diversity of domains. We confirmed that this generative data augmentation method improved the mean average precision by up to 28.65% for the BBox detection task and the average precision by up to 13.73% for the keypoint detection task for vineyard shoot detection. Our generative data augmentation method D4 is expected to simultaneously solve the cost and domain diversity issues of training data generation in agriculture and improve the generalization performance of detection models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。