arXiv:2504.10568cs.CV2025-04NeurIPS被引 15

构建农业多模态理解基准,提升模型在真实场景下的知识推理能力

AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark

  • 基于真实农技对话构建多模态数据集,避免人工构造偏差
  • 包含746道选择题与746道开放题,覆盖五大关键农业主题
  • 开源数据助力小模型性能提升,推动农业AI可信决策研究

我们提出AgMMU,一个面向农业知识密集型场景的挑战性真实世界基准,用于评估和推进视觉-语言模型(VLMs)的发展。不同于依赖众包提示的现有数据集,AgMMU源自116,231条普通农户与美国农业部授权合作推广专家的真实对话。通过三阶段流程:自动化知识提取、问答生成与人工验证,构建了(i)包含746道多选题(MCQs)和746道开放题(OEQs)的评测集,以及(ii)包含57,079个跨模态事实的开发语料库AgBase,涵盖昆虫识别、物种识别、病害分类、症状描述和管理建议五大高风险农业议题。对12个主流VLMs的测评显示,模型在细粒度感知与事实对齐方面存在显著差距;开源模型性能远落后于专有模型。仅在AgBase上进行简单微调,即可使开源模型在复杂开放题上的平均表现提升11.6%,缩小差距,并激励未来在知识提取与蒸馏策略上的创新。我们期望AgMMU能推动特定领域知识融合与可信赖农业AI决策的研究。

原文摘要 · Abstract (English)

We present AgMMU, a challenging real-world benchmark for evaluating and advancing vision-language models (VLMs) in the knowledge-intensive domain of agriculture. Unlike prior datasets that rely on crowdsourced prompts, AgMMU is distilled from 116,231 authentic dialogues between everyday growers and USDA-authorized Cooperative Extension experts. Through a three-stage pipeline: automated knowledge extraction, QA generation, and human verification, we construct (i) AgMMU, an evaluation set of 746 multiple-choice questions (MCQs) and 746 open-ended questions (OEQs), and (ii) AgBase, a development corpus of 57,079 multimodal facts covering five high-stakes agricultural topics: insect identification, species identification, disease categorization, symptom description, and management instruction. Benchmarking 12 leading VLMs reveals pronounced gaps in fine-grained perception and factual grounding. Open-sourced models trail after proprietary ones by a wide margin. Simple fine-tuning on AgBase boosts open-sourced model performance on challenging OEQs for up to 11.6% on average, narrowing this gap and also motivating future research to propose better strategies in knowledge extraction and distillation from AgBase. We hope AgMMU stimulates research on domain-specific knowledge integration and trustworthy decision support in agriculture AI development.

农业AI多模态知识图谱评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。