arXiv:2512.15977cs.CV2025-12被引 2

测试27个农业图像数据集,发现视觉语言模型零样本表现远逊于专用模型。

Are vision-language models ready to zero-shot replace supervised classification models in agriculture?

  • 用零样本和多选提示测试27个农业数据集,评估多种视觉语言模型性能。
  • 最佳模型Gemini-3 Pro平均准确率仅62%,开放提示下普遍低于25%。
  • 提示设计与评估方法显著影响结果,适合有领域知识的辅助诊断场景。

视觉语言模型(VLMs)被广泛视为通用视觉识别解决方案,但其在农业决策支持中的可靠性仍不明确。本研究在来自AgML集合的27个农业图像分类数据集上评估了多种开源与闭源VLMs,涵盖162类、24.8万张图像,涉及植物病害、虫害与损伤、以及作物与杂草种类识别。所有任务中,零样本VLMs显著低于专用监督基线模型YOLO11的性能。在多选提示下,表现最佳的VLM(Gemini-3 Pro)平均准确率为62%,而开放提示下的原始准确率通常低于25%。采用基于LLM的语义判断可提升开放提示准确率(如从~21%升至~30%),并改变模型排名,表明评估方法对结论有显著影响。在开源模型中,Qwen-VL-72B表现最优,在受限提示下接近闭源模型,但仍落后于顶尖专有系统。任务层面分析显示,作物与杂草分类比虫害与损伤识别更易,后者始终是最具挑战性的类别。总体而言,当前即用型VLMs尚不适合作为独立农业诊断系统,但在结合受限接口、显式标签本体和领域感知评估策略时,可作为辅助组件使用。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are increasingly proposed as general-purpose solutions for visual recognition tasks, yet their reliability for agricultural decision support remains poorly understood. We benchmark a diverse set of open-source and closed-source VLMs on 27 agricultural image classification datasets from the AgML collection (https://github.com/Project-AgML), spanning 162 classes and 248,000 images across plant disease, pest and damage, and plant and weed species identification. Across all tasks, zero-shot VLMs substantially underperform a supervised task-specific baseline (YOLO11), which consistently achieves markedly higher accuracy than any foundation model. Under multiple-choice prompting, the best-performing VLM (Gemini-3 Pro) reaches approximately 62% average accuracy, while open-ended prompting yields much lower performance, with raw accuracies typically below 25%. Applying LLM-based semantic judging increases open-ended accuracy (e.g., from ~21% to ~30% for top models) and alters model rankings, demonstrating that evaluation methodology meaningfully affects reported conclusions. Among open-source models, Qwen-VL-72B performs best, approaching closed-source performance under constrained prompting but still trailing top proprietary systems. Task-level analysis shows that plant and weed species classification is consistently easier than pest and damage identification, which remains the most challenging category across models. Overall, these results indicate that current off-the-shelf VLMs are not yet suitable as standalone agricultural diagnostic systems, but can function as assistive components when paired with constrained interfaces, explicit label ontologies, and domain-aware evaluation strategies.

视觉语言模型农业图像零样本学习评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。