arXiv:2410.01407cs.CV2024-10被引 40

为农业畜牧业定制的视觉语言模型,提升细粒度识别准确率。

AgriCLIP: Adapting CLIP for Agriculture and Livestock via Domain-Specialized Cross-Model Alignment

  • 用自定义提示生成海量图文对,构建60万条农业专用数据集
  • 融合对比与自监督学习,同时捕捉全局语义和局部细节特征
  • 在20项下游任务中零样本准确率提升7.8%,适合农业智能识别

基于大规模图像-文本数据,大规模视觉语言预训练展现出出色的零样本能力,已应用于多个领域。然而,基于通用网络爬取数据训练的模型在特定领域表现不佳,可能源于领域偏移。已有研究通过构建专用图像-文本数据集解决了部分领域(如医疗)的问题,但可持续农业与畜牧业仍缺乏大规模专用数据集。此外,该领域需细粒度特征学习,因下游任务差异细微(如营养缺乏检测、畜禽品种分类)。为此,我们提出AgriCLIP,一个专用于农业与畜牧业的视觉语言基础模型。首先,提出大规模数据集ALive,采用定制化提示生成策略克服专家标注稀缺问题。ALive覆盖作物、畜禽及渔业,包含约60万张图像-文本对。其次,设计融合对比学习与自监督学习的训练流程,以同时学习全局语义与局部细粒度领域特异性特征。在20个多样化下游任务上的实验表明,AgriCLIP框架有效,相较于标准CLIP通过专用数据集微调,在平均零样本分类准确率上提升7.8%。ALive数据集与代码可于GitHub获取。

原文摘要 · Abstract (English)

Capitalizing on vast amount of image-text data, large-scale vision-language pre-training has demonstrated remarkable zero-shot capabilities and has been utilized in several applications. However, models trained on general everyday web-crawled data often exhibit sub-optimal performance for specialized domains, likely due to domain shift. Recent works have tackled this problem for some domains (e.g., healthcare) by constructing domain-specialized image-text data. However, constructing a dedicated large-scale image-text dataset for sustainable area of agriculture and livestock is still open to research. Further, this domain desires fine-grained feature learning due to the subtle nature of the downstream tasks (e.g, nutrient deficiency detection, livestock breed classification). To address this we present AgriCLIP, a vision-language foundational model dedicated to the domain of agriculture and livestock. First, we propose a large-scale dataset, named ALive, that leverages customized prompt generation strategy to overcome the scarcity of expert annotations. Our ALive dataset covers crops, livestock, and fishery, with around 600,000 image-text pairs. Second, we propose a training pipeline that integrates both contrastive and self-supervised learning to learn both global semantic and local fine-grained domain-specialized features. Experiments on diverse set of 20 downstream tasks demonstrate the effectiveness of AgriCLIP framework, achieving an absolute gain of 7.8\% in terms of average zero-shot classification accuracy, over the standard CLIP adaptation via domain-specialized ALive dataset. Our ALive dataset and code can be accessible at \href{https://github.com/umair1221/AgriCLIP/tree/main}{Github}.

农业视觉多模态细粒度识别数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。