arXiv:2411.00988cs.CV2024-11EMNLP被引 2

用网络文本增强图像表征,让零样本图像分类在数据稀少场景下表现更优

Retrieval-enriched zero-shot image classification in low-resource domains

  • 通过检索网页文本信息,补充查询图像和类别原型的语义
  • 在医学、罕见植物等低资源领域上显著超越现有方法
  • 无需训练,适合数据稀缺但需快速部署的场景

低资源领域因数据与标注稀少,给视觉理解任务带来巨大挑战,尤其在视觉语言模型(VLM)的应用中尚未充分探索。尽管当前VLM在高资源领域表现良好,但在预训练数据中覆盖极少(如每类仅少数图像)的低资源概念上仍表现不佳。本文提出一种全新的零样本低资源图像分类方法——CoRE(Retrieval-based Retrieval Enrichment),采用无需训练的检索策略,从大规模网络爬取数据库中检索相关文本信息,用于增强查询图像与类别原型的表征。该方法通过引入更广泛的上下文信息,显著提升分类性能。我们在一个新构建的涵盖医疗影像、稀有植物和电路图等多样低资源领域的基准上验证了该方法,实验表明其优于依赖合成数据生成或模型微调的现有最先进方法。

原文摘要 · Abstract (English)

Low-resource domains, characterized by scarce data and annotations, present significant challenges for language and visual understanding tasks, with the latter much under-explored in the literature. Recent advancements in Vision-Language Models (VLM) have shown promising results in high-resource domains but fall short in low-resource concepts that are under-represented (e.g. only a handful of images per category) in the pre-training set. We tackle the challenging task of zero-shot low-resource image classification from a novel perspective. By leveraging a retrieval-based strategy, we achieve this in a training-free fashion. Specifically, our method, named CoRE (Combination of Retrieval Enrichment), enriches the representation of both query images and class prototypes by retrieving relevant textual information from large web-crawled databases. This retrieval-based enrichment significantly boosts classification performance by incorporating the broader contextual information relevant to the specific class. We validate our method on a newly established benchmark covering diverse low-resource domains, including medical imaging, rare plants, and circuits. Our experiments demonstrate that CORE outperforms existing state-of-the-art methods that rely on synthetic data generation and model fine-tuning.

零样本分类低资源学习视觉语言模型信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。