arXiv:2507.21742cs.CV2025-07ICCV被引 2

让图像检索摆脱类别依赖,提升对新类别的泛化能力

Adversarial Reconstruction Feedback for Robust Fine-grained Generalization

  • 用对抗性重构反馈机制,分离类别特定与通用差异特征
  • 在多个数据集上超越现有方法,显著提升未见类别的检索准确率
  • 适合需要跨类别泛化的细粒度图像检索场景

现有细粒度图像检索(FGIR)方法主要依赖预定义类别的监督信号来学习判别性表示,但会无意引入类别特定语义,导致对未见类别的泛化能力受限。为此,我们提出AdvRF——一种新颖的对抗性重构反馈框架,旨在学习类别无关的差异表示。具体地,AdvRF将FGIR重构成视觉差异重建任务,通过融合检索模型的类别感知差异定位与重建模型的类别无关特征学习。重建模型揭示检索模型忽略的残差差异,迫使检索模型提升定位精度;而检索模型优化后的信号则引导重建模型增强重建能力。最终,检索模型实现差异定位,重建模型编码出类别无关表示,并通过知识蒸馏传递给检索模型以实现高效部署。定量与定性评估表明,AdvRF在多个主流细粒度与粗粒度数据集上均取得优异性能。

原文摘要 · Abstract (English)

Existing fine-grained image retrieval (FGIR) methods predominantly rely on supervision from predefined categories to learn discriminative representations for retrieving fine-grained objects. However, they inadvertently introduce category-specific semantics into the retrieval representation, creating semantic dependencies on predefined classes that critically hinder generalization to unseen categories. To tackle this, we propose AdvRF, a novel adversarial reconstruction feedback framework aimed at learning category-agnostic discrepancy representations. Specifically, AdvRF reformulates FGIR as a visual discrepancy reconstruction task via synergizing category-aware discrepancy localization from retrieval models with category-agnostic feature learning from reconstruction models. The reconstruction model exposes residual discrepancies overlooked by the retrieval model, forcing it to improve localization accuracy, while the refined signals from the retrieval model guide the reconstruction model to improve its reconstruction ability. Consequently, the retrieval model localizes visual differences, while the reconstruction model encodes these differences into category-agnostic representations. This representation is then transferred to the retrieval model through knowledge distillation for efficient deployment. Quantitative and qualitative evaluations demonstrate that our AdvRF achieves impressive performance on both widely-used fine-grained and coarse-grained datasets.

细粒度检索对抗学习泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。