构建番茄病害多模态理解数据集,评测视觉模型诊断能力
TomaMMU: A Comprehensive Multimodal Understanding Benchmark for Tomato Leaf Diseases

- 构建三阶段数据流水线生成2.8万张图像与21万条问答对
- 14个主流模型在细粒度识别和推理上均表现不足,挑战题准确率低于90%
- 微调后模型在难题上准确率达96.09%,适合农业视觉研究者使用
为填补该空白,我们提出TomaMMU,一个大规模番茄叶片病害多模态理解数据集,并构建TomaBench基准,用于评估视觉语言模型(VLMs)在番茄病害理解上的表现。TomaMMU包含28,808张高质量图像,覆盖15类病害,以及213,119条人工标注的视觉问答对,通过数据收集、人工标注和问答生成三阶段流程构建。基于此,TomaBench将七个农业任务组织成从基础感知到病理理解再到专家诊断的三级层次化分类体系,系统评估模型从低层视觉识别到高层诊断推理的能力。任务涵盖症状识别、分类关系与诊断推理,全面考察模型对植物病理学的理解程度。14个先进VLMs在细粒度识别和事实性推理方面存在明显短板,挑战性选择题与开放问答表现均不理想。结果表明当前VLM难以将视觉感知转化为可靠诊断知识,亟需领域适配。在TomaMMU上进行简单微调显著缩小差距,挑战题准确率提升至96.09%,优于近期多数VLM,指明未来研究方向。所有数据与代码已公开于https://huggingface.co/datasets/enalis/TomaMMU。
原文摘要 · Abstract (English)
To address this gap, we introduce TomaMMU, a large-scale Tomato leaf disease MultiModal Understanding dataset, alongside TomaBench, a benchmark for evaluating VLMs on tomato disease understanding. TomaMMU comprises 28,808 high-quality images spanning 15 categories and 213,119 human-annotated visual question-answer pairs, generated through a three-stage pipeline comprising Data Collection, Human Annotation, and Question-Answer Generation. Building on this foundation, TomaBench organizes seven agricultural tasks into a hierarchical three-level taxonomy spanning Basic Perception, Pathology Understanding, and Expert Diagnosis, which together enable systematic evaluation from low-level visual recognition to high-level diagnostic reasoning. The tasks assess visual symptom recognition, taxonomic relationships, and diagnostic reasoning, offering a comprehensive view of how well models grasp plant pathology. Our results pronounced gaps in fine-grained recognition and factually grounded reasoning with 14 state-of-the-art VLMs, consistently underperforming on both challenging MCQs and open-ended questions. These results suggest that current VLMs struggle to translate visual perception into reliable diagnostic knowledge, motivating the need for targeted domain adaptation. Simple fine-tuning on TomaMMU substantially narrows this gap, boosting accuracy on challenging MCQs to 96.09%, outperforming recent VLMs, and pointing toward promising directions for future work. All data and code is available in https://huggingface.co/datasets/enalis/TomaMMU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。