arXiv:2506.16273cs.CVcs.MM2025-06AAAI被引 2

通过双视觉适配提升细粒度图像检索精度与泛化能力

Fine-grained Image Retrieval via Dual-Vision Adaptation

  • 输入样本与特征协同调整,不修改预训练模型参数
  • 在3个分布内和3个分布外数据集上均表现优异
  • 参数量少,适合资源受限场景的细粒度检索应用

细粒度图像检索(FGIR)面临学习判别性视觉表征的挑战。现有主流方法通常采用两种范式:在语义嵌入空间施加成对相似性约束,或引入定位子网络微调整个模型。但这些方法易过拟合训练数据,遗忘大规模预训练所得知识,导致泛化能力下降。本文提出双视觉适配(DVA)方法,通过协同样本与特征适应,引导冻结的预训练模型完成FGIR。具体地,设计对象感知适配,修改输入样本以增强模型对关键物体及内部元素的感知;同时提出上下文适配,引入少量可学习参数进行特征调整,使任务更贴近预训练目标。为平衡效率与性能,进一步提出判别感知迁移,利用知识蒸馏将对象感知适配中的判别知识传递至图像编码器。大量实验表明,DVA参数量少,在三个分布内和三个分布外的细粒度数据集上均取得良好效果。

原文摘要 · Abstract (English)

Fine-Grained Image Retrieval~(FGIR) faces challenges in learning discriminative visual representations to retrieve images with similar fine-grained features. Current leading FGIR solutions typically follow two regimes: enforce pairwise similarity constraints in the semantic embedding space, or incorporate a localization sub-network to fine-tune the entire model. However, such two regimes tend to overfit the training data while forgetting the knowledge gained from large-scale pre-training, thus reducing their generalization ability. In this paper, we propose a Dual-Vision Adaptation (DVA) approach for FGIR, which guides the frozen pre-trained model to perform FGIR through collaborative sample and feature adaptation. Specifically, we design Object-Perceptual Adaptation, which modifies input samples to help the pre-trained model perceive critical objects and elements within objects that are helpful for category prediction. Meanwhile, we propose In-Context Adaptation, which introduces a small set of parameters for feature adaptation without modifying the pre-trained parameters. This makes the FGIR task using these adjusted features closer to the task solved during the pre-training. Additionally, to balance retrieval efficiency and performance, we propose Discrimination Perception Transfer to transfer the discriminative knowledge in the object-perceptual adaptation to the image encoder using the knowledge distillation mechanism. Extensive experiments show that DVA has fewer learnable parameters and performs well on three in-distribution and three out-of-distribution fine-grained datasets.

细粒度检索视觉适配知识蒸馏预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。