arXiv:2504.07718cs.CV2025-04被引 15

用多模态参考提升细粒度图文检索精度

Multi-modal Reference Learning for Fine-grained Text-to-Image Retrieval

  • 构建图文联合参考,融合同一对象的视觉与文本细节
  • 在RSTPReid上达56.2%的Rank1准确率,超现有方法5.6%
  • 适合需要精准图文匹配的图像检索场景

细粒度文本到图像检索旨在根据文本查询找到特定目标图像。现有方法通常假设每张训练图像的文本描述是准确的,但文本描述常存在歧义,难以捕捉图像中的判别性视觉细节,导致表征学习不准确。为缓解文本歧义影响,我们提出多模态参考学习框架,通过多模态参考构建模块将同一对象的全部视觉与文本细节整合为综合参考,从而辅助后续表征学习与检索相似性计算。具体地,设计参考引导的表征学习模块,利用多模态参考学习更精确的视觉与文本表征;并引入基于参考的优化方法,通过对象参考计算参考相似性,对初始检索结果进行精炼。在五个细粒度文本到图像检索数据集上进行了广泛实验,所提方法在多种任务中均优于当前最优方法。例如,在文本到人物图像检索数据集RSTPReid上,达到56.2%的Rank1准确率,超越近期方法CFine 5.6%。

原文摘要 · Abstract (English)

Fine-grained text-to-image retrieval aims to retrieve a fine-grained target image with a given text query. Existing methods typically assume that each training image is accurately depicted by its textual descriptions. However, textual descriptions can be ambiguous and fail to depict discriminative visual details in images, leading to inaccurate representation learning. To alleviate the effects of text ambiguity, we propose a Multi-Modal Reference learning framework to learn robust representations. We first propose a multi-modal reference construction module to aggregate all visual and textual details of the same object into a comprehensive multi-modal reference. The multi-modal reference hence facilitates the subsequent representation learning and retrieval similarity computation. Specifically, a reference-guided representation learning module is proposed to use multi-modal references to learn more accurate visual and textual representations. Additionally, we introduce a reference-based refinement method that employs the object references to compute a reference-based similarity that refines the initial retrieval results. Extensive experiments are conducted on five fine-grained text-to-image retrieval datasets for different text-to-image retrieval tasks. The proposed method has achieved superior performance over state-of-the-art methods. For instance, on the text-to-person image retrieval dataset RSTPReid, our method achieves the Rank1 accuracy of 56.2\%, surpassing the recent CFine by 5.6\%.

图文检索多模态学习细粒度匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。