arXiv:2604.07814cs.CV2026-04

用专家验证的推理链训练农业视觉语言模型,提升病害诊断准确率与可解释性。

AgriChain Visually Grounded Expert Verified Reasoning for Interpretable Agricultural Vision Language Models

  • 构建1.1万张专家标注叶部图像数据集,含疾病标签、置信度与验证后的推理链条。
  • 模型在1000张测试图像上达73.1%准确率,显著优于Gemini、GPT-4o等主流模型。
  • 生成解释紧扣关键视觉特征,适合农业专家与开发者用于可信农业AI部署。

真实农业中,精准且可解释的植物病害诊断仍是视觉语言模型的重大挑战。本文提出AgriChain,一个约1.1万张由专家筛选的叶片图像数据集,涵盖多种作物与病害,每张图像均配有(i)病害标签,(ii)校准的置信度评分(高/中/低),以及(iii)经专业农工验证的链式思考(CoT)推理说明。初始解释由GPT-4o生成,再由农业工程师使用标准化描述符(如病斑颜色、边缘、分布)进行审核。我们在AgriChain上微调Qwen2.5-VL-3B,得到专用模型AgriChain-VL3B,实现疾病联合预测与视觉引导的推理生成。在1000张图像的测试集上,该模型达到73.1%的Top-1准确率(宏F1=0.466;加权F1=0.655),优于Gemini 1.5 Flash、Gemini 2.5 Pro和GPT-4o Mini等强基线模型。生成的解释与专家推理高度一致,持续引用关键视觉线索。结果表明,专家验证的推理监督显著提升准确性与可解释性,弥合通用多模态模型与人类专长之间的差距,推动可信赖、全球可用的可持续农业人工智能发展。数据集与代码已公开于:https://github.com/hazzanabeel12-netizen/agrichain

原文摘要 · Abstract (English)

Accurate and interpretable plant disease diagnosis remains a major challenge for vision-language models (VLMs) in real-world agriculture. We introduce AgriChain, a dataset of approximately 11,000 expert-curated leaf images spanning diverse crops and pathologies, each paired with (i) a disease label, (ii) a calibrated confidence score (High/Medium/Low), and (iii) an expert-verified chain-of-thought (CoT) rationale. Draft explanations were first generated by GPT-4o and then verified by a professional agricultural engineer using standardized descriptors (e.g., lesion color, margin, and distribution). We fine-tune Qwen2.5-VL-3B on AgriChain, resulting in a specialized model termed AgriChain-VL3B, to jointly predict diseases and generate visually grounded reasoning. On a 1,000-image test set, our CoT-supervised model achieves 73.1% top-1 accuracy (macro F1 = 0.466; weighted F1 = 0.655), outperforming strong baselines including Gemini 1.5 Flash, Gemini 2.5 Pro, and GPT-4o Mini. The generated explanations align closely with expert reasoning, consistently referencing key visual cues. These findings demonstrate that expert-verified reasoning supervision significantly enhances both accuracy and interpretability, bridging the gap between generic multimodal models and human expertise, and advancing trustworthy, globally deployable AI for sustainable agriculture. The dataset and code are publicly available at: https://github.com/hazzanabeel12-netizen/agrichain

农业AI视觉语言模型可解释性病害诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。