arXiv:2508.17117cs.CVcs.AI2025-08被引 7

构建植物科学视觉问答数据集,推动农业决策的智能诊断。

PlantExpertVQA: A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science

  • 基于45个开源数据集合成76万条带图像的问答对,覆盖38种作物和89种病害。
  • 当前主流多模态模型在该数据集上表现差,但小样本微调即显著提升性能。
  • 适合研究农业视觉语言模型、可解释性诊断与领域适配的学者使用。

现有植物病害数据集主要面向分类与检测任务,难以支持视觉-语言模型进行交互式、推理型诊断。为此,我们提出PlantExpertVQA,一个大规模视觉问答(VQA)数据集,旨在推动农业决策中视觉-语言模型的发展。该数据集整合自45个开源数据集,包括广泛使用的PlantVillage,包含765,186条高质量问答对,覆盖150,841张图像,涉及38种作物和89种病害。问题按认知复杂度分为3个层级、9个类别,由专家指导设计,并通过两阶段自动化流程生成:先基于图像元数据模板合成问答,再经多阶段语言重构优化。所有问答经领域专家迭代审核以确保科学准确性与相关性。我们发现,当前前沿视觉-语言模型(包括近期开源指令微调的多模态大模型)在PlantExpertVQA上表现不佳。然而,仅用少量数据对一个20亿参数的紧凑模型进行参数高效微调,即可在所有问题类别上实现显著性能提升,证明了其在领域适应中的有效性。

原文摘要 · Abstract (English)

Existing plant-disease datasets target classification and detection, leaving vision-language models unable to support interactive, reasoning-based diagnosis. To address this, we present PlantExpertVQA, a large-scale visual question answering (VQA) dataset designed to advance vision-language models for agricultural decision-making. It is compiled from 45 open-source datasets, including the widely used PlantVillage corpus, and comprises 765,186 high-quality question-answer (QA) pairs grounded over 150,841 images spanning 38 crop species and 89 disease conditions. Questions are organized into 3 levels of cognitive complexity and 9 distinct categories. Each was phrased following expert guidance and generated via an automated two-stage pipeline: template-based QA synthesis from image metadata, followed by multi-stage linguistic re-engineering. The dataset was iteratively reviewed by domain experts for scientific accuracy and relevance. We find that current frontier vision-language models, including recent open-source instruction-tuned multimodal LLMs, perform poorly on PlantExpertVQA. However, parameter-efficient fine-tuning of a compact 2B-parameter model on a small fraction of the dataset yields substantial improvements across all question categories, demonstrating its effectiveness for domain adaptation.

视觉问答农业AI多模态模型领域适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。