arXiv:2507.20519cs.CV2025-07ICCV被引 31

农业视觉语言模型新基准,专为精准农事识别设计

AgroBench: Vision-Language Model Benchmark in Agriculture

  • 由农学家标注的7大农业主题评测集
  • 覆盖203种作物与682类病害,评估模型细粒度识别能力
  • 发现开源模型在杂草识别中近乎随机,适合农业AI研究者

精准的自动化农业任务理解,如病害识别,对可持续作物生产至关重要。近年来,视觉语言模型(VLMs)通过文本交互有望拓展农业应用范围。本文提出AgroBench(农艺师AI基准),一个涵盖七个农业主题的评测基准,覆盖农业工程关键领域并贴近实际耕作场景。不同于现有农业VLM基准,AgroBench由专家农学家标注。其包含203种作物类别和682种病害类别,全面评估VLM性能。在该基准上的评估显示,当前VLM在细粒度识别任务中仍有提升空间。值得注意的是,在杂草识别任务中,多数开源VLM表现接近随机水平。基于广泛的主题和专家标注数据,我们分析了模型错误类型,并提出未来VLM发展的潜在路径。数据集与代码已公开于 https://dahlian00.github.io/AgroBenchPage/。

原文摘要 · Abstract (English)

Precise automated understanding of agricultural tasks such as disease identification is essential for sustainable crop production. Recent advances in vision-language models (VLMs) are expected to further expand the range of agricultural tasks by facilitating human-model interaction through easy, text-based communication. Here, we introduce AgroBench (Agronomist AI Benchmark), a benchmark for evaluating VLM models across seven agricultural topics, covering key areas in agricultural engineering and relevant to real-world farming. Unlike recent agricultural VLM benchmarks, AgroBench is annotated by expert agronomists. Our AgroBench covers a state-of-the-art range of categories, including 203 crop categories and 682 disease categories, to thoroughly evaluate VLM capabilities. In our evaluation on AgroBench, we reveal that VLMs have room for improvement in fine-grained identification tasks. Notably, in weed identification, most open-source VLMs perform close to random. With our wide range of topics and expert-annotated categories, we analyze the types of errors made by VLMs and suggest potential pathways for future VLM development. Our dataset and code are available at https://dahlian00.github.io/AgroBenchPage/ .

农业AI视觉语言模型病害识别基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。