用自一致性提升农业图像诊断准确率
Self-Consistency in Vision-Language Models for Precision Agriculture: Multi-Response Consensus for Crop Disease Management
- 通过提示工程让模型模拟植物病理专家评估结果
- 多回复投票选出最一致的病害诊断,准确率提至87.8%
- 轻量设计适合手机端实时使用,适合田间场景
精准农业依赖图像分析进行病害识别与治疗建议,但现有视觉语言模型在农业领域表现不佳。本文提出一种面向农业的框架,结合基于提示的专家评估与自一致性机制,提升模型可靠性。创新点包括:(1) 设计提示协议,将语言模型配置为可扩展的植物病理专家以评估图像分析输出;(2) 采用余弦一致性自投票机制,从农业图像生成多个候选回答,并利用领域适配嵌入选择语义最一致的诊断。在微调后的PaliGemma模型上应用于玉米叶片病害识别,相比标准贪婪解码,诊断准确率从82.2%提升至87.8%,症状分析从38.9%提升至52.2%,治疗建议从27.8%提升至43.3%。系统保持轻量化,可在移动设备部署,支持资源受限环境下的实时农业决策。结果表明该方法在复杂田间条件下具有显著应用潜力。
原文摘要 · Abstract (English)
Precision agriculture relies heavily on accurate image analysis for crop disease identification and treatment recommendation, yet existing vision-language models (VLMs) often underperform in specialized agricultural domains. This work presents a domain-aware framework for agricultural image processing that combines prompt-based expert evaluation with self-consistency mechanisms to enhance VLM reliability in precision agriculture applications. We introduce two key innovations: (1) a prompt-based evaluation protocol that configures a language model as an expert plant pathologist for scalable assessment of image analysis outputs, and (2) a cosine-consistency self-voting mechanism that generates multiple candidate responses from agricultural images and selects the most semantically coherent diagnosis using domain-adapted embeddings. Applied to maize leaf disease identification from field images using a fine-tuned PaliGemma model, our approach improves diagnostic accuracy from 82.2\% to 87.8\%, symptom analysis from 38.9\% to 52.2\%, and treatment recommendation from 27.8\% to 43.3\% compared to standard greedy decoding. The system remains compact enough for deployment on mobile devices, supporting real-time agricultural decision-making in resource-constrained environments. These results demonstrate significant potential for AI-driven precision agriculture tools that can operate reliably in diverse field conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。