构建农业工具增强型多模态智能体评测基准,提升精准决策能力。
AgroTools: A Benchmark for Tool-Augmented Multimodal Agents in Agriculture

- 设计包含14种农具的可执行环境,评估模型使用工具的全过程。
- 9个开源与4个闭源模型在工具规划与执行恢复上表现不佳。
- 适合研究高精度农业智能体的学者和开发者使用。
农业决策日益依赖能将视觉观测转化为可靠可执行动作的多模态系统。然而,现有农业多模态基准主要评估最终答案正确性,缺乏对模型利用外部工具完成精密工作流的支持。本文提出AgroTools,一个面向农业领域工具增强型多模态智能体的评测基准。该基准包含539个问答实例及1,097张异构农业图像,涵盖五个任务类别,并提供由14种农业工具构成的可执行环境。每个查询均带有结构化工具使用轨迹,支持过程级执行质量与结果级任务成功率的双重评估。我们在AgroTools上对9个开源和4个闭源多模态大模型进行了评测,结果显示当前模型在农业工具使用场景中仍远未可靠,存在工具规划、参数生成、执行恢复和最终答案合成等明显瓶颈。我们希望AgroTools能推动高精度农业应用中多模态智能体的研究。基准与评估代码已开放:https://huggingface.co/datasets/AgroTools/AgroTools。
原文摘要 · Abstract (English)
Agricultural decision-making increasingly requires multimodal systems that can transform visual observations into reliable, executable actions. However, existing agricultural multimodal benchmarks mainly evaluate final-answer correctness and provide limited support for assessing whether models can use external tools to complete precision-sensitive workflows. In this paper, we introduce AgroTools, a benchmark for evaluating tool-augmented multimodal agents in agriculture. AgroTools contains 539 question-answer instances paired with 1,097 heterogeneous agricultural images, spanning five task families and an executable environment of 14 agricultural tools. Each query is annotated with structured tool-use traces, enabling a dual-view evaluation of both process-level execution quality and outcome-level task success. We benchmark 9 open-source and 4 closed-source multimodal large language models on AgroTools. Results show that current models remain far from reliable in agricultural tool-use settings, with clear bottlenecks in tool planning, argument generation, execution recovery, and final-answer synthesis. We hope AgroTools will support future research on multimodal agents for high-precision agricultural applications. The benchmark and evaluation are available at https://huggingface.co/datasets/AgroTools/AgroTools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。