用视觉数据构建农业对话模型,让AI懂行农事。
AgroGPT: Efficient Agricultural Vision-Language Model with Expert Tuning
- 用农业图像数据+大模型生成指令数据,构造70K专家训练集
- 训练出能精准识别细粒度农业概念的高效多模态模型
- 适合农业研究者、智能农技助手开发者使用
大型多模态对话模型虽取得显著进展,但普遍存在领域差距问题,难以在新领域开展复杂对话。现有方法依赖特定领域的图文数据进行指令微调,但农业等领域缺乏此类数据。本文提出一种新方法:利用多样化的农业视觉数据,提取类别信息,并通过大语言模型生成专家级指令数据,构建了70k规模的农业指令数据集AgroInstruct。基于此,我们训练出AgroGPT——一个高效的农业多模态语言模型,可进行复杂农业对话并提供专业见解。我们还开发了AgroEvals评估基准,对比结果表明,AgroGPT在细粒度农业概念识别和多模态农业问答中表现优异,具备农业专家能力。代码、数据集与模型已开源。
原文摘要 · Abstract (English)
Significant progress has been made in advancing large multimodal conversational models (LMMs), capitalizing on vast repositories of image-text data available online. Despite this progress, these models often encounter substantial domain gaps, hindering their ability to engage in complex conversations across new domains. Recent efforts have aimed to mitigate this issue, albeit relying on domain-specific image-text data to curate instruction-tuning data. However, many domains, such as agriculture, lack such vision-language data. In this work, we propose an approach to construct instruction-tuning data that harnesses vision-only data for the agriculture domain. We utilize diverse agricultural datasets spanning multiple domains, curate class-specific information, and employ large language models (LLMs) to construct an expert-tuning set, resulting in a 70k expert-tuning dataset called AgroInstruct. Subsequently, we expert-tuned and created AgroGPT, an efficient LMM that can hold complex agriculture-related conversations and provide useful insights. We also develop AgroEvals for evaluation and compare {AgroGPT's} performance with large open and closed-source models. {AgroGPT} excels at identifying fine-grained agricultural concepts, can act as an agriculture expert, and provides helpful information for multimodal agriculture questions. The code, datasets, and models are available at https://github.com/awaisrauf/agroGPT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。