评测大模型对农业遥感的理解能力,发现其在空间推理上仍有短板。
Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind
- 构建涵盖13类任务的农业遥感评测集AgroMind
- 20个开源与4个闭源模型测试,人类表现落后于部分模型
- 揭示大模型在细粒度识别和空间推理中的不足
大型多模态模型(LMMs)在多个领域表现出色,但针对农业遥感(RS)的综合性评测基准仍十分缺乏。现有农业遥感评测数据集存在场景多样性不足、任务设计过于简单等问题。为此,我们提出AgroMind,一个覆盖空间感知、物体理解、场景理解与场景推理四个维度的农业遥感基准,包含13种任务类型,从作物识别、健康监测到环境分析。通过整合8个公开数据集与1个私有农田地块数据集,构建了包含27,247个问答对和19,615张图像的高质量评估集。数据处理流程包括多源数据采集、格式标准化与标注优化,并基于系统化任务定义生成多样化的农业相关问题。采用LMMs进行推理并开展细致分析。我们评估了20个开源及4个闭源模型,实验显示模型在空间推理和细粒度识别方面存在显著性能差距,值得注意的是,人类表现落后于若干领先的大模型。AgroMind为农业遥感建立了标准化评估框架,揭示了大模型在领域知识上的局限性,指出了未来研究的关键挑战。数据与代码可访问:https://rssysu.github.io/AgroMind/
原文摘要 · Abstract (English)
Large Multimodal Models (LMMs) has demonstrated capabilities across various domains, but comprehensive benchmarks for agricultural remote sensing (RS) remain scarce. Existing benchmarks designed for agricultural RS scenarios exhibit notable limitations, primarily in terms of insufficient scene diversity in the dataset and oversimplified task design. To bridge this gap, we introduce AgroMind, a comprehensive agricultural remote sensing benchmark covering four task dimensions: spatial perception, object understanding, scene understanding, and scene reasoning, with a total of 13 task types, ranging from crop identification and health monitoring to environmental analysis. We curate a high-quality evaluation set by integrating eight public datasets and one private farmland plot dataset, containing 27,247 QA pairs and 19,615 images. The pipeline begins with multi-source data pre-processing, including collection, format standardization, and annotation refinement. We then generate a diverse set of agriculturally relevant questions through the systematic definition of tasks. Finally, we employ LMMs for inference, generating responses, and performing detailed examinations. We evaluated 20 open-source LMMs and 4 closed-source models on AgroMind. Experiments reveal significant performance gaps, particularly in spatial reasoning and fine-grained recognition, it is notable that human performance lags behind several leading LMMs. By establishing a standardized evaluation framework for agricultural RS, AgroMind reveals the limitations of LMMs in domain knowledge and highlights critical challenges for future work. Data and code can be accessed at https://rssysu.github.io/AgroMind/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。