arXiv:2412.00139cs.CV2024-12

通过动态微调提升文本到图像检索在开放域下的鲁棒性

EFSA: Episodic Few-Shot Adaptation for Text-to-Image Retrieval

  • 测试时基于检索结果和合成描述微调模型,实现快速适应
  • 在超百万图像的开放域中跨8个不同视觉领域显著提升性能
  • 适合需要高泛化能力的开放域图像检索场景

文本到图像检索对管理多样化视觉内容至关重要,但现有基准多依赖小规模、单领域数据集,难以反映真实世界复杂性。预训练视觉语言模型在处理简单负样本时表现良好,但在开放域场景下面对视觉相似却错误的难负样本时表现不佳。为此,我们提出一种新的测试时自适应框架——情景式少样本适应(EFSA),通过在查询对应的 top-k 检索候选及其生成的合成描述上进行微调,动态适配预训练模型以应对当前查询的领域特征。在来自八个差异显著的视觉领域及包含超过一百万张图像的开放域检索池上的评估表明,EFSA 在保持模型泛化能力的同时,显著提升了跨域检索性能。本工作展示了情景式少样本适应在关键且研究不足的开放域文本到图像检索任务中的潜力。

原文摘要 · Abstract (English)

Text-to-image retrieval is a critical task for managing diverse visual content, but common benchmarks for the task rely on small, single-domain datasets that fail to capture real-world complexity. Pre-trained vision-language models tend to perform well with easy negatives but struggle with hard negatives--visually similar yet incorrect images--especially in open-domain scenarios. To address this, we introduce Episodic Few-Shot Adaptation (EFSA), a novel test-time framework that adapts pre-trained models dynamically to a query's domain by fine-tuning on top-k retrieved candidates and synthetic captions generated for them. EFSA improves performance across diverse domains while preserving generalization, as shown in evaluations on queries from eight highly distinct visual domains and an open-domain retrieval pool of over one million images. Our work highlights the potential of episodic few-shot adaptation to enhance robustness in the critical and understudied task of open-domain text-to-image retrieval.

文本到图像少样本学习检索开放域

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。