为农业大模型微调提供可复现的框架与评估标准
Fine-Tuning General-Purpose Large Language Models for Agricultural Applications:A Reproducible Framework and Evaluation Protocol Based on Qwen3-8B
- 基于Qwen3-8B构建农业专用微调流程,融合数据治理与安全控制
- 设计涵盖病虫害咨询、农事管理等四类任务的评估协议
- 强调专家评审与事实一致性,适合农业AI研究者参考
通用大语言模型在开放域问答、信息抽取和文本生成方面表现优异,但农业应用具有领域性强、地域依赖、时效敏感和高风险的特点。若缺乏数据治理、专家评估与证据约束,农业助手可能给出不可靠的作物病害、农药使用、施肥或政策解读建议。本文不报告未经实际训练与专家评估的性能结论,而是提出AgriTune-R框架,推荐使用公开可验证的Qwen3-8B作为基础模型,整合农业数据治理、指令构建、LoRA/QLoRA参数高效微调、检索增强生成、专家评估与高风险问题安全控制。贡献包括:(1)农业LLM适配的结构化工作流;(2)农业知识问答、病虫害咨询、种植管理与政策解释的评估协议;(3)融合事实性、安全性、证据一致性和不确定性表达的专家评审量表;(4)明确区分协议设计与实证结论,为未来研究提供可执行基准。
原文摘要 · Abstract (English)
General-purpose large language models (LLMs) have demonstrated strong abilities in opendomain question answering, information extraction, and text generation. Agricultural applications, however, are domain-specific, region-dependent, time-sensitive, and safety-critical. Without data governance, expert evaluation, and evidence constraints, an agricultural assistant mayproduce unreliable advice on crop diseases, pesticide use, fertilization, or policy interpretation.To avoid presenting unverified simulated numbers as real experimental findings, this paper doesnot report any model-performance claims that have not been produced by an actual training runand expert evaluation. Instead, we propose AgriTune-R, a reproducible and auditable frameworkfor adapting general-purpose LLMs to agricultural tasks. The framework selects the publiclyverifiable Qwen3-8B model as the recommended base model and integrates agricultural datagovernance, instruction construction, LoRA/QLoRA parameter-efficient fine-tuning, retrievalaugmented generation, expert evaluation, and safety control for high-risk questions. The contributions are: (1) a structured workflow for agricultural LLM adaptation; (2) an evaluationprotocol for agricultural knowledge QA, pest and disease consultation, cultivation management,and policy explanation; (3) an expert-review rubric combining factuality, safety, evidence consistency, and uncertainty expression; and (4) a clear separation between protocol design andempirical conclusions, providing an executable baseline for future empirical studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。