让图像检索精准锁定具体物体,而非仅匹配语义
Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval
- 引入边界框锚定目标对象,确保检索实例一致
- 在16万组数据上验证,显著提升实例保真度
- 适合需要精确物体匹配的视觉搜索场景
组合图像检索(CIR)通过结合参考图像与修改文本实现灵活的多模态查询,但其本质上侧重语义匹配,难以在不同上下文中可靠检索指定实例。实践中,保持具体实例的准确性往往比宽泛语义更重要。为此,本文提出对象锚定的组合图像检索(OACIR),要求严格保证实例级一致性。为推动该任务研究,我们构建了首个大规模、跨领域基准OACIRR,包含超过16万组四元组及四个含困难负样本的候选图库。每个四元组附加参考图像中的边界框,精确定位目标对象,确保实例保留。针对OACIR任务,我们提出AdaFocal框架,包含上下文感知注意力调制器,可自适应增强指定实例区域的关注度,动态平衡锚定实例与整体语义上下文之间的焦点。大量实验表明,AdaFocal显著优于现有组合检索模型,尤其在保持实例级保真度方面表现突出,为该挑战性任务建立了强基线,并开启更灵活、实例感知检索系统的新方向。
原文摘要 · Abstract (English)
Composed Image Retrieval (CIR) has demonstrated significant potential by enabling flexible multimodal queries that combine a reference image and modification text. However, CIR inherently prioritizes semantic matching, struggling to reliably retrieve a user-specified instance across contexts. In practice, emphasizing concrete instance fidelity over broad semantics is often more consequential. In this work, we propose Object-Anchored Composed Image Retrieval (OACIR), a novel fine-grained retrieval task that mandates strict instance-level consistency. To advance research on this task, we construct OACIRR (OACIR on Real-world images), the first large-scale, multi-domain benchmark comprising over 160K quadruples and four challenging candidate galleries enriched with hard-negative instance distractors. Each quadruple augments the compositional query with a bounding box that visually anchors the object in the reference image, providing a precise and flexible way to ensure instance preservation. To address the OACIR task, we propose AdaFocal, a framework featuring a Context-Aware Attention Modulator that adaptively intensifies attention within the specified instance region, dynamically balancing focus between the anchored instance and the broader compositional context. Extensive experiments demonstrate that AdaFocal substantially outperforms existing compositional retrieval models, particularly in maintaining instance-level fidelity, thereby establishing a robust baseline for this challenging task while opening new directions for more flexible, instance-aware retrieval systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。