用大模型融合视觉专家判断,实现可解释的植物病害诊断。
Fusing Perceptual Vision Experts with Multimodal Large Language Models for Explainable Plant Disease Diagnosis: From Benchmark Imagery to Real-World Robotic Field Validation

- 双视觉模型+大语言模型协同决策,用结构化证据生成诊断结论。
- 在真实农田数据上准确率达98.9%~99.3%,冲突情况下提升7.6个百分点。
- 输出诊断理由、风险等级和治疗紧迫性,适合农业机器人部署使用。
精准田间植物病害诊断需融合不确定且矛盾的感知证据。我们提出混合分层多智能体框架(H²MAF),结合EfficientNet-B3与ConvNeXt-Tiny在决策层的融合,以及开放权重多模态大语言模型(MLLMs)Gemma 4 E4B与Qwen3.5 4B进行语义仲裁,通过结构化JSON证据生成可解释的诊断结果、风险等级、治疗紧迫性及财务影响。该框架在14,364张图像(1,370张测试集)上评估,涵盖PlantDoc(2,922张,27类)及两个非公开、由康奈尔机器人持续采集的田间数据集:Stage 2(20 GB;4,215张)和Stage 4(40 GB;7,227张),覆盖早疫病、晚疫病和褐斑病等,在非受控田间条件下表现优异。在PlantDoc上,Gemma将准确率从63.9%提升至68.5%,在41.7%的CNN冲突子集上提升7.6个百分点。康奈尔数据集准确率达99.3%与98.9%,分歧仅1.7%-4.1%,体现MLLM在冲突场景下的有效性。Gemma的高风险错误为0.14-0.5点,而Qwen则存在3.5-14.4点的过度预警。结果表明,MLLM仲裁是可解释农业AI与机器人田间决策支持的有前景方向,但依赖校准。
原文摘要 · Abstract (English)
Accurate field plant disease diagnosis requires reliable fusion of uncertain and conflicting perceptual evidence. We present the Hybrid Hierarchical Multi-Agent Framework (H$^{2}$MAF), combining decision-level fusion of EfficientNet-B3 and ConvNeXt-Tiny with semantic arbitration by open-weight multimodal large language models (MLLMs), Gemma 4 E4B and Qwen3.5 4B, using structured JSON evidence to generate explainable diagnoses, risk levels, treatment urgency, and financial exposure. (H$^{2}$MAF) is evaluated on 14,364 images (1,370 test images) across PlantDoc (2,922 images, 27 classes) and two non-public, continuously captured Cornell robot-acquired field datasets: Stage 2 (20 GB; 4,215 images) and Stage 4 (40 GB; 7,227 images), covering Early Blight, Late Blight, and Septoria Leaf Spot under uncontrolled field conditions. On PlantDoc, Gemma improves accuracy from 63.9% to 68.5%, achieving +7.6 points on the 41.7% CNN-conflict subset. Cornell accuracies reach 99.3% and 98.9%, with only 1.7-4.1% disagreement, demonstrating conflict-dependent MLLM utility. The critical-risk error of gemma is 0.14-0.5 points, whereas Qwen overflags by 3.5-14.4 points. These results establish MLLM arbitration as a promising, yet calibration-dependent, approach for explainable agricultural AI and robotic field decision support. Github Link: https://github.com/Applied-AI-Research-Lab/Explainable-AI-Plant-Disease-Detection
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。