arXiv:2507.21773cs.CL2025-07AAAI被引 5

首个中文农业大模型评测基准,覆盖6大类29子领域

AgriEval: A Comprehensive Chinese Agricultural Benchmark for Large Language Models

  • 构建涵盖6大类29子领域的农业知识评估体系
  • 包含1.47万道选择题与2167道开放问答题,规模领先
  • 适合农业AI研究者、教育科技开发者使用

在农业领域,大语言模型(LLMs)的部署受限于训练数据和评估基准的缺乏。为此,我们提出AgriEval,首个全面的中文农业基准,具有三大特点:(1) 全面能力评估。涵盖农业六大类别与29个子类别,覆盖记忆、理解、推理和生成四类核心认知场景;(2) 高质量数据。数据源自高校考试与作业,自然真实,可评估模型应用知识与专家决策能力;(3) 多样格式与大规模。包含14,697道多选题与2,167道开放式问答题,是当前最广泛的农业基准。我们在51个开源与商业大模型上进行实验,结果表明多数模型准确率不足60%,凸显农业LLM的巨大发展潜力。此外,我们深入分析影响性能的因素并提出优化策略。AgriEval已公开于https://github.com/YanPioneer/AgriEval/。

原文摘要 · Abstract (English)

In the agricultural domain, the deployment of large language models (LLMs) is hindered by the lack of training data and evaluation benchmarks. To mitigate this issue, we propose AgriEval, the first comprehensive Chinese agricultural benchmark with three main characteristics: (1) Comprehensive Capability Evaluation. AgriEval covers six major agriculture categories and 29 subcategories within agriculture, addressing four core cognitive scenarios: memorization, understanding, inference, and generation. (2) High-Quality Data. The dataset is curated from university-level examinations and assignments, providing a natural and robust benchmark for assessing the capacity of LLMs to apply knowledge and make expert-like decisions. (3) Diverse Formats and Extensive Scale. AgriEval comprises 14,697 multiple-choice questions and 2,167 open-ended question-and-answer questions, establishing it as the most extensive agricultural benchmark available to date. We also present comprehensive experimental results over 51 open-source and commercial LLMs. The experimental results reveal that most existing LLMs struggle to achieve 60% accuracy, underscoring the developmental potential in agricultural LLMs. Additionally, we conduct extensive experiments to investigate factors influencing model performance and propose strategies for enhancement. AgriEval is available at https://github.com/YanPioneer/AgriEval/.

农业AI大模型评测中文数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。