不依赖视觉信息的图像检索新方法,提升复杂组合查询准确率
Towards Vision-Free CIR: Attribute-Augmented Scoring and LLM-Based Reranking for Zero-Shot Composed Image Retrieval

- 用属性增强文本匹配弥补图像信息损失
- 用大模型重排序验证语义一致性,提升检索精度
- 适合零样本组合图像检索场景,尤其关注语义与细节平衡
近期研究证明,将图像表示为文本的‘无视觉’方法在标准图像检索中有效。但其能否应对更复杂的组合图像检索(CIR)任务仍不明确,因文本描述存在固有信息损失。本文提出一种无视觉CIR框架,通过两项关键技术解决该问题:(1) 属性增强的混合打分,通过显式属性匹配补偿丢失的视觉细节;(2) 基于大语言模型的重排序,验证候选结果的语义一致性。在开放域CIRR数据集上的实验表明,本方法优于现有零样本CIR方法(R@1达44.04%,提升8.79%)。在FashionIQ上,结果揭示了语义推理与细粒度视觉匹配间的权衡。消融实验证明,属性增强打分与大模型重排序均能持续提升性能。
原文摘要 · Abstract (English)
Recent work has shown that "Vision-Free'' approaches (representing images as text) can be effective for standard image retrieval tasks. However, it remains unclear whether this paradigm can effectively handle a more complex, multimodal task, Composed Image Retrieval (CIR), due to the inherent information loss in textual descriptions. In this paper, we introduce a Vision-Free CIR framework that addresses this challenge through two key techniques: (1) Attribute-Augmented Hybrid Scoring, which compensates for lost visual details via explicit attribute matching, and (2) LLM-Based Reranking, which verifies semantic consistency of top candidates. Experiments on the open-domain CIRR dataset show that our approach outperforms existing Zero-shot CIR methods (44.04% R@1, +8.79%). On FashionIQ, our results highlight the trade-off between semantic reasoning and fine-grained visual matching. Ablation studies reveal that both attribute-augmented scoring and LLM-Based Reranking consistently improve performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。