arXiv:2508.08632cs.AI2025-08被引 24

专为农业打造的LLM生态,解决领域模型缺失问题

AgriGPT: a Large Language Model Ecosystem for Agriculture

  • 构建多智能体数据引擎,生成34.2万条农业高质量问答数据
  • 引入三通道检索增强生成,推理准确率显著提升
  • 开源完整生态,适合农业研究者与基层从业者使用

尽管大型语言模型发展迅速,但其在农业领域的应用仍受限于缺乏领域专用模型、标注数据集和可靠评估框架。为此,我们提出AgriGPT——一个面向农业应用的领域专用大模型生态系统。核心是设计了一个可扩展的多智能体数据引擎,系统性地整合可信数据源,构建了包含34.2万条样本的高质量标准化问答数据集Agri-342K。基于该数据集训练的AgriGPT支持从农户到政策制定者的广泛用户群体。为增强事实准确性,采用三通道检索增强生成(Tri-RAG)框架,融合密集检索、稀疏检索与多跳知识图谱推理,显著提升模型推理可靠性。针对全面评估,提出AgriBench-13K基准套件,涵盖13项不同类型与复杂度的任务。实验表明,AgriGPT在领域适配与推理能力上均显著优于通用大模型。除模型外,AgriGPT还构建了结构化的数据构建、检索增强生成与领域评估一体化的模块化生态系统,为科学与产业专用大模型开发提供可复用框架。所有模型、数据集与代码将公开发布,赋能农业社区,尤其惠及欠发达地区,推动开放且具影响力的科研。

原文摘要 · Abstract (English)

Despite the rapid progress of Large Language Models (LLMs), their application in agriculture remains limited due to the lack of domain-specific models, curated datasets, and robust evaluation frameworks. To address these challenges, we propose AgriGPT, a domain-specialized LLM ecosystem for agricultural usage. At its core, we design a multi-agent scalable data engine that systematically compiles credible data sources into Agri-342K, a high-quality, standardized question-answer (QA) dataset. Trained on this dataset, AgriGPT supports a broad range of agricultural stakeholders, from practitioners to policy-makers. To enhance factual grounding, we employ Tri-RAG, a three-channel Retrieval-Augmented Generation framework combining dense retrieval, sparse retrieval, and multi-hop knowledge graph reasoning, thereby improving the LLM's reasoning reliability. For comprehensive evaluation, we introduce AgriBench-13K, a benchmark suite comprising 13 tasks with varying types and complexities. Experiments demonstrate that AgriGPT significantly outperforms general-purpose LLMs on both domain adaptation and reasoning. Beyond the model itself, AgriGPT represents a modular and extensible LLM ecosystem for agriculture, comprising structured data construction, retrieval-enhanced generation, and domain-specific evaluation. This work provides a generalizable framework for developing scientific and industry-specialized LLMs. All models, datasets, and code will be released to empower agricultural communities, especially in underserved regions, and to promote open, impactful research.

农业AI大模型数据构建检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。