构建首个植物病害多模态数据集与评测基准,推动农业AI诊断发展。
LeafNet: A Large-Scale Dataset and Comprehensive Benchmark for Foundational Vision-Language Understanding of Plant Diseases
- 构建18.6万张叶片图像与1.4万组问答对,覆盖97类病害。
- 验证12个先进模型性能,细粒度识别准确率不足65%。
- 揭示多模态模型显著优于纯视觉模型,适合农业领域研究者使用。
基础模型与视觉语言预训练显著推进了视觉语言模型(VLMs)的发展,实现了视觉与语言数据的多模态处理。然而,在植物病理学等特定农业任务中,由于缺乏大规模、全面的多模态图像-文本数据集和评测基准,其应用仍受限。为此,我们提出了LeafNet多模态数据集和LeafBench视觉问答评测基准,用于系统评估VLMs在植物病害理解方面的能力。数据集包含186,000张叶片数字图像,涵盖97种病害类别,配有元数据,并生成13,950组跨六类关键农业任务的问答对。问题涵盖症状识别、分类关系及诊断推理等维度。在LeafBench上对12个前沿VLMs进行评测,结果显示模型性能在任务间差异显著:健康/病害二分类准确率超90%,而细粒度病原体与物种识别仍低于65%。对比视觉模型与VLMs表明,多模态架构具有明显优势:微调后的VLMs显著优于传统视觉模型,证明融合语言表征可大幅提升诊断精度。这些发现揭示了当前VLMs在植物病理学应用中的关键差距,强调了LeafBench作为方法改进与进展评估的严谨框架的重要性。代码已开源:https://github.com/EnalisUs/LeafBench。
原文摘要 · Abstract (English)
Foundation models and vision-language pre-training have significantly advanced Vision-Language Models (VLMs), enabling multimodal processing of visual and linguistic data. However, their application in domain-specific agricultural tasks, such as plant pathology, remains limited due to the lack of large-scale, comprehensive multimodal image--text datasets and benchmarks. To address this gap, we introduce LeafNet, a comprehensive multimodal dataset, and LeafBench, a visual question-answering benchmark developed to systematically evaluate the capabilities of VLMs in understanding plant diseases. The dataset comprises 186,000 leaf digital images spanning 97 disease classes, paired with metadata, generating 13,950 question-answer pairs spanning six critical agricultural tasks. The questions assess various aspects of plant pathology understanding, including visual symptom recognition, taxonomic relationships, and diagnostic reasoning. Benchmarking 12 state-of-the-art VLMs on our LeafBench dataset, we reveal substantial disparity in their disease understanding capabilities. Our study shows performance varies markedly across tasks: binary healthy--diseased classification exceeds 90% accuracy, while fine-grained pathogen and species identification remains below 65%. Direct comparison between vision-only models and VLMs demonstrates the critical advantage of multimodal architectures: fine-tuned VLMs outperform traditional vision models, confirming that integrating linguistic representations significantly enhances diagnostic precision. These findings highlight critical gaps in current VLMs for plant pathology applications and underscore the need for LeafBench as a rigorous framework for methodological advancement and progress evaluation toward reliable AI-assisted plant disease diagnosis. Code is available at https://github.com/EnalisUs/LeafBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。