arXiv:2510.21757cs.CV2025-10

用轻量级方法提升农业病害识别模型的可靠性,适合资源有限地区使用。

Agro-Consensus: Semantic Self-Consistency in Vision-Language Models for Crop Disease Management in Developing Countries

  • 通过语义聚类和相似度筛选,从多个生成结果中选出最一致的诊断描述。
  • 在800张图像上,10个候选生成下准确率达83.1%,优于基线77.5%。
  • 支持人工确认作物类型,适配网络差、专家少的发展中国家场景。

印度、肯尼亚和尼日利亚等发展中国家的农业病害管理面临专家匮乏、网络不稳定和成本限制等挑战,难以部署大规模AI系统。本文提出一种低成本自一致性框架,提升视觉语言模型(VLM)在农作物图像描述中的可靠性。该方法采用轻量级(80MB)预训练嵌入模型进行语义聚类,通过余弦相似度筛选出包含诊断、症状、分析、治疗与预防建议的最连贯描述。引入人机协同组件,用户确认作物种类以过滤错误生成,提高输入质量。基于微调后的3B参数PaliGemma模型,在公开的PlantVillage数据集上评估,每张图像生成最多21个候选结果。单聚类共识方法在10个候选生成下达到83.1%的峰值准确率,优于贪婪解码的77.5%。当在前四个聚类中任一包含正确答案时,准确率提升至94.0%,超过基线方法的88.5%。

原文摘要 · Abstract (English)

Agricultural disease management in developing countries such as India, Kenya, and Nigeria faces significant challenges due to limited access to expert plant pathologists, unreliable internet connectivity, and cost constraints that hinder the deployment of large-scale AI systems. This work introduces a cost-effective self-consistency framework to improve vision-language model (VLM) reliability for agricultural image captioning. The proposed method employs semantic clustering, using a lightweight (80MB) pre-trained embedding model to group multiple candidate responses. It then selects the most coherent caption -- containing a diagnosis, symptoms, analysis, treatment, and prevention recommendations -- through a cosine similarity-based consensus. A practical human-in-the-loop (HITL) component is incorporated, wherein user confirmation of the crop type filters erroneous generations, ensuring higher-quality input for the consensus mechanism. Applied to the publicly available PlantVillage dataset using a fine-tuned 3B-parameter PaliGemma model, our framework demonstrates improvements over standard decoding methods. Evaluated on 800 crop disease images with up to 21 generations per image, our single-cluster consensus method achieves a peak accuracy of 83.1% with 10 candidate generations, compared to the 77.5% baseline accuracy of greedy decoding. The framework's effectiveness is further demonstrated when considering multiple clusters; accuracy rises to 94.0% when a correct response is found within any of the top four candidate clusters, outperforming the 88.5% achieved by a top-4 selection from the baseline.

农业AI视觉语言模型自一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。