提升图像检索中细粒度属性修改的精准度
DQE-CIR: Distinctive Query Embeddings through Learnable Attribute Weights and Target Relative Negative Sampling in Composed Image Retrieval
- 通过可学习属性权重强化文本引导下的视觉特征
- 设计目标相关负样本采样,减少语义混淆
- 适合需要精细图像修改检索的研究与应用
组合图像检索(CIR)旨在通过参考图像和描述修改意图的文本共同定位目标图像。现有方法多基于对比学习,将真实图像视为唯一正例,其余均为负例,导致相关但非目标图像被错误排斥(相关性抑制),不同修改意图在嵌入空间中重叠(语义混淆),使查询表示缺乏区分性,尤其在细粒度属性修改时表现不佳。为此,我们提出DQE-CIR:通过可学习属性权重与目标相对负采样,显式建模目标相关性以学习更具区分性的查询嵌入。该方法引入可学习属性权重,根据修改文本强调差异性视觉特征,实现语言与视觉更精确对齐;同时提出目标相对负采样,构建目标相关相似度分布,从中间区域选取有信息量的负样本,排除易分负例与模糊假负例。该策略显著提升细粒度属性修改的检索可靠性,增强查询区分性,缓解语义相似但无关候选带来的混淆。
原文摘要 · Abstract (English)
Composed image retrieval (CIR) addresses the task of retrieving a target image by jointly interpreting a reference image and a modification text that specifies the intended change. Most existing methods are still built upon contrastive learning frameworks that treat the ground truth image as the only positive instance and all remaining images as negatives. This strategy inevitably introduces relevance suppression, where semantically related yet valid images are incorrectly pushed away, and semantic confusion, where different modification intents collapse into overlapping regions of the embedding space. As a result, the learned query representations often lack discriminativeness, particularly at fine-grained attribute modifications. To overcome these limitations, we propose distinctive query embeddings through learnable attribute weights and target relative negative sampling (DQE-CIR), a method designed to learn distinctive query embeddings by explicitly modeling target relative relevance during training. DQE-CIR incorporates learnable attribute weighting to emphasize distinctive visual features conditioned on the modification text, enabling more precise feature alignment between language and vision. Furthermore, we introduce target relative negative sampling, which constructs a target relative similarity distribution and selects informative negatives from a mid-zone region that excludes both easy negatives and ambiguous false negatives. This strategy enables more reliable retrieval for fine-grained attribute changes by improving query discriminativeness and reducing confusion caused by semantically similar but irrelevant candidates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。