首个农业病虫害多模态大模型,助力精准识别与防治
Agri-LLaVA: Knowledge-Infused Large Multimodal Assistant on Agricultural Pests and Diseases
- 构建首个农业病虫害多模态指令数据集,含40万条数据
- 提出知识注入训练法,使模型在病虫害图像理解上表现更优
- 适合农业科研、植保人员及AI研究者使用
通用领域的大规模多模态模型虽有显著进展,但在农业等专业领域的应用仍面临挑战。农业作为全球经济支柱,正面临病虫害复杂多变、传播迅速、抗性强等问题。本文首次构建了农业领域多模态指令跟随数据集,涵盖221种病虫害,共约40万条数据,旨在探索病虫害防控中的独特难题。基于此数据集,提出知识注入训练方法,开发出农业多模态对话系统Agri-LLaVA。为推动该领域发展,设计了多样且具有挑战性的评估基准。实验表明,Agri-LLaVA在农业多模态对话和视觉理解任务中表现优异,提供新思路与工具。项目已开源,资源详见https://github.com/Kki2Eve/Agri-LLaVA。
原文摘要 · Abstract (English)
In the general domain, large multimodal models (LMMs) have achieved significant advancements, yet challenges persist in applying them to specific fields, especially agriculture. As the backbone of the global economy, agriculture confronts numerous challenges, with pests and diseases being particularly concerning due to their complexity, variability, rapid spread, and high resistance. This paper specifically addresses these issues. We construct the first multimodal instruction-following dataset in the agricultural domain, covering over 221 types of pests and diseases with approximately 400,000 data entries. This dataset aims to explore and address the unique challenges in pest and disease control. Based on this dataset, we propose a knowledge-infused training method to develop Agri-LLaVA, an agricultural multimodal conversation system. To accelerate progress in this field and inspire more researchers to engage, we design a diverse and challenging evaluation benchmark for agricultural pests and diseases. Experimental results demonstrate that Agri-LLaVA excels in agricultural multimodal conversation and visual understanding, providing new insights and approaches to address agricultural pests and diseases. By open-sourcing our dataset and model, we aim to promote research and development in LMMs within the agricultural domain and make significant contributions to tackle the challenges of agricultural pests and diseases. All resources can be found at https://github.com/Kki2Eve/Agri-LLaVA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。