arXiv:2604.05583cs.CV2026-04

针对图像检索过拟合问题,提出权重正则微调方法提升泛化能力。

WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval

  • 用反梯度方向的对抗扰动正则化模型权重,增强训练难度
  • 在有限三元组数据下,显著缩小不同模型/数据集间的泛化差距
  • 适合低资源场景下的视觉语言图像检索任务

组合图像检索(CIR)旨在根据参考图像和修改文本检索目标图像。现有方法主要依赖微调视觉-语言预训练模型,但我们发现这些方法普遍存在严重过拟合问题,尤其在三元组数据有限时表现不佳。为此,我们对基于视觉-语言模型的CIR中的过拟合现象进行了系统性研究,揭示了不同模型与数据集间存在显著且此前被忽视的泛化差距。受此启发,我们提出WRF4CIR——一种用于CIR的权重正则化微调网络。具体而言,在微调过程中,我们对模型权重施加对抗扰动,扰动方向与梯度下降相反。直观上,该方法增加了拟合训练数据的难度,从而有效缓解了在有限三元组监督下的过拟合问题。大量基准数据集上的实验表明,WRF4CIR显著缩小了泛化差距,并在性能上优于现有方法。

原文摘要 · Abstract (English)

Composed Image Retrieval (CIR) task aims to retrieve target images based on reference images and modification texts. Current CIR methods primarily rely on fine-tuning vision-language pre-trained models. However, we find that these approaches commonly suffer from severe overfitting, posing challenges for CIR with limited triplet data. To better understand this issue, we present a systematic study of overfitting in VLP-based CIR, revealing a significant and previously overlooked generalization gap across different models and datasets. Motivated by these findings, we introduce WRF4CIR, a Weight-Regularized Fine-tuning network for CIR. Specifically, during the fine-tuning process, we apply adversarial perturbations to the model weights for regularization, where these perturbations are generated in the opposite direction of gradient descent. Intuitively, WRF4CIR increases the difficulty of fitting the training data, which helps mitigate overfitting in CIR under limited triplet supervision. Extensive experiments on benchmark datasets demonstrate that WRF4CIR significantly narrows the generalization gap and achieves substantial improvements over existing methods.

图像检索微调正则视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。