arXiv:2604.20358cs.CV2026-04中稿 · CVPR被引 16

解决图像检索中的噪声标注问题,提升模型鲁棒性。

ConeSep: Cone-based Robust Noise-Unlearning Compositional Network for Composed Image Retrieval

论文配图:ConeSep: Cone-based Robust Noise-Unlearning Compositional Network for Composed Image Retrieval
图 1 · 摘自论文原文
  • 基于锥形几何构建噪声边界,精确定位错误标注
  • 引入对角负例学习,增强语义区分能力
  • 通过最优传输机制避免去噪副作用,适合高噪声场景

组成图像检索(CIR)任务通过参考图像和修改文本实现灵活检索,但严重依赖昂贵且易错的三元组标注。本文系统研究了由标注引入的噪声三元组对应(NTC)问题。发现硬噪声(即参考与目标图像高度相似但修改文本错误)对现有噪声对应学习(NCL)方法构成独特挑战,因其破坏传统“小损失假设”。我们识别并阐明三大被忽视的关键挑战:(C1) 模态抑制、(C2) 负样本缺失、(C3) 去噪反噬。为此,提出锥形鲁棒去噪组合网络(ConeSep)。首先提出几何保真度量化,理论建立并实践估计噪声边界以精准定位噪声对应;其次引入负边界学习,为每个查询显式学习嵌入空间中的“对角负组合”作为语义相反锚点;最后设计基于边界的定向去噪,将噪声修正过程建模为最优传输问题,优雅规避去噪反噬。在基准数据集FashionIQ和CIRR上的大量实验表明,ConeSep显著优于当前最先进方法,充分验证了其有效性和鲁棒性。

原文摘要 · Abstract (English)

The Composed Image Retrieval (CIR) task provides a flexible retrieval paradigm via a reference image and modification text, but it heavily relies on expensive and error-prone triplet annotations. This paper systematically investigates the Noisy Triplet Correspondence (NTC) problem introduced by annotations. We find that NTC noise, particularly ``hard noise'' (i.e., the reference and target images are highly similar but the modification text is incorrect), poses a unique challenge to existing Noise Correspondence Learning (NCL) methods because it breaks the traditional ``small loss hypothesis''. We identify and elucidate three key, yet overlooked, challenges in the NTC task, namely (C1) Modality Suppression, (C2) Negative Anchor Deficiency, and (C3) Unlearning Backlash. To address these challenges, we propose a Cone-based robuSt noisE-unlearning comPositional network (ConeSep). Specifically, we first propose Geometric Fidelity Quantization, theoretically establishing and practically estimating a noise boundary to precisely locate noisy correspondence. Next, we introduce Negative Boundary Learning, which learns a ``diagonal negative combination'' for each query as its explicit semantic opposite-anchor in the embedding space. Finally, we design Boundary-based Targeted Unlearning, which models the noisy correction process as an optimal transport problem, elegantly avoiding Unlearning Backlash. Extensive experiments on benchmark datasets (FashionIQ and CIRR) demonstrate that ConeSep significantly outperforms current state-of-the-art methods, which fully demonstrates the effectiveness and robustness of our method.

图像检索去噪学习最优传输鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。