arXiv:2412.18806cs.CVcs.IR2024-12被引 2

微调模型实现精准开放词汇图像检索,准确率显著提升。

FOR: Finetuning for Object Level Open Vocabulary Image Retrieval

  • 用封闭标签微调模型,保留视觉语言关联性
  • 在三个数据集上最高提升8 mAP@50点
  • 适合小样本半监督场景,标注成本低

随着大规模数据集的普及,通过开放词汇文本查询准确检索目标物体图像的任务日益重要。当前主流方法直接使用预训练的CLIP模型,不进行领域适配,通过后处理平衡精度与效率。本文提出FOR:面向对象级开放词汇图像检索的微调方法,可在仅使用封闭集标签的情况下对目标数据集进行微调,同时保持开放词汇检索所需的视觉-语言关联。FOR基于两个设计:针对任务定制的CLIP头部变体,以及多目标训练框架的集成。这两项设计使准确率显著提升,在三个数据集上相比最先进方法最高提升8 mAP@50点。此外,我们证明了FOR在半监督设置下也有效,即使仅有少量数据标注也能取得优异效果。

原文摘要 · Abstract (English)

As working with large datasets becomes standard, the task of accurately retrieving images containing objects of interest by an open set textual query gains practical importance. The current leading approach utilizes a pre-trained CLIP model without any adaptation to the target domain, balancing accuracy and efficiency through additional post-processing. In this work, we propose FOR: Finetuning for Object-centric Open-vocabulary Image Retrieval, which allows finetuning on a target dataset using closed-set labels while keeping the visual-language association crucial for open vocabulary retrieval. FOR is based on two design elements: a specialized decoder variant of the CLIP head customized for the intended task, and its coupling within a multi-objective training framework. Together, these design choices result in a significant increase in accuracy, showcasing improvements of up to 8 mAP@50 points over SoTA across three datasets. Additionally, we demonstrate that FOR is also effective in a semi-supervised setting, achieving impressive results even when only a small portion of the dataset is labeled.

图像检索开放词汇微调半监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。