arXiv:2505.19259cs.LGcs.AI2025-05被引 4

用专门数据训练小模型,让AI更好做农业决策。

Towards Large Reasoning Models for Agriculture

  • 构建农业推理专用数据集与评测基准
  • 小模型在消费级显卡上达到36%准确率
  • 适合农业科研与智能决策系统开发者

农业决策涉及复杂且情境依赖的推理,作物选择、农事措施和干预策略高度依赖地理、气候与经济条件。传统大语言模型因推理能力有限,难以胜任此类任务。我们提出大型推理模型(LRMs)更适用于此类结构化领域推理。为此,我们引入AgReason,首个专家标注的开放式农业推理基准,包含100个问题。对十三个开源与专有模型的评估显示,LRMs表现优于传统模型,但仍有挑战,最强的Gemini基线仅达36%准确率。我们还构建了44.6万条问答对的数据集AgThoughts,含人工审核与合成推理链。基于此,我们开发了可在消费级GPU运行的AgThinker小推理模型套件,证明该数据集能有效激活大模型的农业推理能力。

原文摘要 · Abstract (English)

Agricultural decision-making involves complex, context-specific reasoning, where choices about crops, practices, and interventions depend heavily on geographic, climatic, and economic conditions. Traditional large language models (LLMs) often fall short in navigating this nuanced problem due to limited reasoning capacity. We hypothesize that recent advances in large reasoning models (LRMs) can better handle such structured, domain-specific inference. To investigate this, we introduce AgReason, the first expert-curated open-ended science benchmark with 100 questions for agricultural reasoning. Evaluations across thirteen open-source and proprietary models reveal that LRMs outperform conventional ones, though notable challenges persist, with the strongest Gemini-based baseline achieving 36% accuracy. We also present AgThoughts, a large-scale dataset of 44.6K question-answer pairs generated with human oversight and equipped with synthetically generated reasoning traces. Using AgThoughts, we develop AgThinker, a suite of small reasoning models that can be run on consumer-grade GPUs, and show that our dataset can be effective in unlocking agricultural reasoning abilities in LLMs. Our project page is here: https://baskargroup.github.io/Ag_reasoning/

农业AI推理模型小模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。