构建农业视觉语言模型推理评估基准,填补复杂场景下的思维链评测空白。
AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture
- 基于思维链(CoT)设计农业多模态问答数据集,聚焦逻辑推理能力
- 包含4759个样本,覆盖零样本场景下的模型推理表现评估
- 揭示30个主流模型在农业任务中的推理短板,推动更智能的农业AI发展
视觉语言模型(VLMs)的最新进展已在多个领域产生重要影响。在农业领域,其多模态能力在精准农业、作物监测、病虫害检测和环境可持续性方面展现出巨大潜力。然而,尽管已有多个视觉问答(VQA)数据集和评估基准被开发,它们往往难以有效评估复杂农业情境下所需的推理与问题解决能力。为弥补这一差距,我们提出AgroCoT,一个融合思维链(CoT)推理的VQA数据集,专门用于评估VLM的推理能力。该数据集包含4,759个精心设计的样本,能够全面、稳健地评估模型在零样本场景下的逻辑推理与问题解决能力。我们对30个代表性VLM进行了评估,涵盖开源与专有模型,结果揭示了当前模型在推理能力上的显著不足,凸显了引入CoT进行评估的重要性。数据集已公开于https://huggingface.co/datasets/AgroCoT/AgroCoT。
原文摘要 · Abstract (English)
Recent advancements in Vision-Language Models (VLMs) have significantly impacted various industries. In agriculture, these multimodal capabilities hold great promise for applications such as precision farming, crop monitoring, pest detection, and environmental sustainability. However, while several Visual Question Answering (VQA) datasets and benchmarks have been developed to assess VLM performance, they often fail to effectively evaluate the critical reasoning and problem-solving skills needed in complex agricultural contexts. To address this gap, we introduce AgroCoT, a VQA dataset that integrates Chain-of-Thought (CoT) reasoning, specifically designed to evaluate the reasoning capabilities of VLMs. With 4,759 carefully curated samples, AgroCoT provides a comprehensive and robust evaluation of reasoning abilities, particularly in zero-shot scenarios, focusing on the models' ability to engage in logical reasoning and effective problem-solving. Our evaluation of 30 representative VLMs, including both proprietary and open-source models, reveals a gap in their reasoning capabilities, which underscores the importance of incorporating CoT for assessments. Our dataset is available at https://huggingface.co/datasets/AgroCoT/AgroCoT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。