arXiv:2607.00374cs.CVcs.AI2026-07中稿 · ECCV

让AI学会真正组合语义,实现零样本图像检索新突破

Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval

论文配图:Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval
图 1 · 摘自论文原文
  • 设计双阶段代理任务,分别聚焦修改内容与上下文补全
  • 在4个基准上达到最优性能,显著提升细粒度语义表达能力
  • 适合研究零样本视觉语言理解与跨模态生成的学者

组合图像检索(CIR)从参考图像和文本修改中检索目标图像。监督式CIR依赖昂贵的三元组数据,而零样本CIR通过基于图文对的代理任务缓解此问题。现有方法主要增强视觉与文本表征,采用预定义的组合机制如伪词注入或线性特征运算,导致组合函数未被学习,限制了对多样化、细粒度语义修改的表达。为此,本文提出FoCo,将组合建模为两个协同阶段:关注修改相关视觉内容,再完成目标语义。通过两个代理任务实现:文本锚定的视觉聚合,以局部文本语义引导视觉内容选择;上下文条件的语义补全,将聚合后的视觉信息结合剩余场景上下文生成连贯组合表示。两者联合训练,并采用跨实例对比目标,促进语义多样性并抑制捷径策略。在四个ZS-CIR基准上的大量实验表明,FoCo表现领先,且泛化能力更强。

原文摘要 · Abstract (English)

Composed Image Retrieval (CIR) retrieves a target image from a reference image and a textual modification. While supervised CIR relies on costly triplets, Zero-Shot CIR (ZS-CIR) alleviates this reliance through proxy tasks trained on image-text pairs. However, existing proxy tasks primarily enhance visual and textual representations to accommodate a predefined composition mechanism such as pseudo-word injection into a frozen text encoder or linear feature arithmetic. As a result, the composition function itself remains unlearned, limiting the model's ability to express diverse and fine-grained semantic modifications. To address this, we propose FoCo, which models composition as two coordinated stages: focusing on modification-relevant visual content, and then completing the target semantics. We realize these through two proxy tasks: text-anchored visual aggregation to selectively gather visual content guided by localized textual semantics, and context-conditioned semantic completion to transform these aggregated visuals with the remaining scene context into a coherent composed representation. The tasks are trained jointly with a cross-instance contrastive objective, encouraging semantic diversity and discouraging shortcut composition strategies. Extensive experiments on four ZS-CIR benchmarks show FoCo's state-of-the-art performance and improved generalization.

图像检索零样本多模态组合生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。