arXiv:2604.09114cs.CVcs.LG2026-04

让图像检索理解穿搭细节变化,提升准确率与可解释性

FIRE-CIR: Fine-grained Reasoning for Composed Fashion Image Retrieval

  • 通过自动生成属性问题,分析参考图与候选图的视觉证据
  • 在Fashion IQ数据集上准确率超越现有方法,支持细粒度修改识别
  • 适合需要精准理解服装修改意图的研究者与应用开发

组合图像检索(CIR)旨在找出被文本描述修改后的目标图像。尽管当前视觉语言模型(VLMs)通过将图像与文本嵌入共享空间实现较好检索性能,但往往无法判断应保留或更改哪些内容,限制了可解释性并导致次优结果,尤其在时尚等细粒度领域表现不佳。本文提出FIRE-CIR,引入组合推理与可解释性机制。该模型不依赖单纯嵌入相似性,而是基于文本生成聚焦属性的视觉问题,并在参考图与候选图中验证对应视觉证据。为训练此推理系统,我们自动构建了一个大规模时尚专用视觉问答数据集,涵盖单图与双图分析任务。检索时,模型利用显式推理对候选结果重新排序,剔除不符合意图的图像。在Fashion IQ基准上的实验表明,FIRE-CIR在检索准确率上优于现有最优方法,同时提供逐属性的可解释决策依据。

原文摘要 · Abstract (English)

Composed image retrieval (CIR) aims to retrieve a target image that depicts a reference image modified by a textual description. While recent vision-language models (VLMs) achieve promising CIR performance by embedding images and text into a shared space for retrieval, they often fail to reason about what to preserve and what to change. This limitation hinders interpretability and yields suboptimal results, particularly in fine-grained domains like fashion. In this paper, we introduce FIRE-CIR, a model that brings compositional reasoning and interpretability to fashion CIR. Instead of relying solely on embedding similarity, FIRE-CIR performs question-driven visual reasoning: it automatically generates attribute-focused visual questions derived from the modification text, and verifies the corresponding visual evidence in both reference and candidate images. To train such a reasoning system, we automatically construct a large-scale fashion-specific visual question answering dataset, containing questions requiring either single- or dual-image analysis. During retrieval, our model leverages this explicit reasoning to re-rank candidate results, filtering out images inconsistent with the intended modifications. Experimental results on the Fashion IQ benchmark show that FIRE-CIR outperforms state-of-the-art methods in retrieval accuracy. It also provides interpretable, attribute-level insights into retrieval decisions.

图像检索视觉推理时尚分析可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。