用视觉大模型增强长尾数据特征,无需语言信息
Enhancing Features in Long-tailed Data Using Large Vision Model
- 仅用视觉大模型提取特征,融合到基础网络的特征图和隐空间
- 在ImageNet-LT和iNaturalist2018上提升长尾分类性能
- 适合无文本标注场景下的长尾识别任务
语言基础模型(如LLM或LVLM)在长尾识别中被广泛研究,但其对语言数据的依赖限制了实际应用。本研究探索使用大型视觉模型(LVM)或视觉基础模型(VFMs)在不依赖任何语言信息的前提下增强长尾数据特征。具体而言,从LVM中提取特征,并将其融合到基线网络的特征图和隐空间中,获得增强特征。此外,在隐空间中设计多种基于原型的损失函数,进一步挖掘增强特征的潜力。实验在两个基准数据集ImageNet-LT和iNaturalist2018上验证了该方法的有效性。
原文摘要 · Abstract (English)
Language-based foundation models, such as large language models (LLMs) or large vision-language models (LVLMs), have been widely studied in long-tailed recognition. However, the need for linguistic data is not applicable to all practical tasks. In this study, we aim to explore using large vision models (LVMs) or visual foundation models (VFMs) to enhance long-tailed data features without any language information. Specifically, we extract features from the LVM and fuse them with features in the baseline network's map and latent space to obtain the augmented features. Moreover, we design several prototype-based losses in the latent space to further exploit the potential of the augmented features. In the experimental section, we validate our approach on two benchmark datasets: ImageNet-LT and iNaturalist2018.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。