针对昆虫视觉理解的通用大模型与数据集,提升农业场景中昆虫识别能力。
Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding
- 构建大规模昆虫多模态数据集,支持视觉-语言联合学习。
- 提出Insect-LLaVA模型,在昆虫问答任务上达到领先水平。
- 采用微特征自监督学习,精准捕捉昆虫细微差异,适合农业研究者使用。
多模态对话生成人工智能在视觉与语言理解方面已展现出强大能力,但现有模型缺乏对昆虫的视觉知识,因其训练数据主要来自通用视觉-语言数据。而昆虫理解是精准农业中的基础问题,有助于推动农业可持续发展。本文提出一种新型多模态对话模型Insect-LLaVA,以促进昆虫领域视觉理解。首先,我们引入一个大规模的多模态昆虫数据集,包含视觉昆虫指令数据,使模型能够学习昆虫的视觉与语义特征。其次,提出Insect-LLaVA,一种面向昆虫视觉理解的通用大语言与视觉助手。为增强昆虫特征学习能力,我们设计了基于局部特征的自监督学习方法,并引入逐块相关注意力机制,以捕捉昆虫图像间的细微差异;同时提出描述一致性损失,通过文本描述进一步优化微特征学习。在新构建的视觉昆虫问答基准测试中,实验结果表明所提方法在昆虫视觉理解任务上表现优异,且在标准昆虫相关任务基准上达到最新技术水平。
原文摘要 · Abstract (English)
Multimodal conversational generative AI has shown impressive capabilities in various vision and language understanding through learning massive text-image data. However, current conversational models still lack knowledge about visual insects since they are often trained on the general knowledge of vision-language data. Meanwhile, understanding insects is a fundamental problem in precision agriculture, helping to promote sustainable development in agriculture. Therefore, this paper proposes a novel multimodal conversational model, Insect-LLaVA, to promote visual understanding in insect-domain knowledge. In particular, we first introduce a new large-scale Multimodal Insect Dataset with Visual Insect Instruction Data that enables the capability of learning the multimodal foundation models. Our proposed dataset enables conversational models to comprehend the visual and semantic features of the insects. Second, we propose a new Insect-LLaVA model, a new general Large Language and Vision Assistant in Visual Insect Understanding. Then, to enhance the capability of learning insect features, we develop an Insect Foundation Model by introducing a new micro-feature self-supervised learning with a Patch-wise Relevant Attention mechanism to capture the subtle differences among insect images. We also present Description Consistency loss to improve micro-feature learning via text descriptions. The experimental results evaluated on our new Visual Insect Question Answering benchmarks illustrate the effective performance of our proposed approach in visual insect understanding and achieve State-of-the-Art performance on standard benchmarks of insect-related tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。