提出INTENT模型,解决图像检索中的噪声问题,提升跨模态匹配鲁棒性。
INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image Retrieval

- 通过频域干预生成视觉不变特征,抑制模态内噪声
- 双目标学习动态调整决策边界,增强正负样本区分能力
- 适用于存在标注错误的真实场景,适合图像检索研究者
组合图像检索(CIR)是一种基于参考图像和修改文本的多模态查询来检索目标图像的挑战性任务。尽管近年取得进展,现有方法假设所有样本均正确匹配,但真实场景中因三元组标注成本高,数据集不可避免包含标注错误,导致不正确三元组。我们提出噪声三元组对应(NTC)问题,并认为噪声可分为跨模态对应噪声与模态内固有噪声:前者源于模态间不匹配,后者来自模态内背景干扰或与粗粒度修改注释无关的视觉因素。然而模态内噪声常被忽视,跨模态噪声研究仍不充分。为此,我们提出视角不变与判别感知噪声缓解网络(INTENT),包含两个组件:视觉不变组合与双目标判别学习。前者通过快速傅里叶变换(FFT)在视觉侧实施因果干预,生成干预后组合特征,强制视觉不变性,使模型在组合时忽略模态内噪声;后者采用正负样本协同优化,构建可扩展决策边界,根据忠诚度动态调整判断,实现稳健对应判别。在两个主流基准数据集上的大量实验表明,INTENT具有优越性和鲁棒性。
原文摘要 · Abstract (English)
Composed Image Retrieval (CIR) is a challenging image retrieval paradigm that enables to retrieve target images based on multimodal queries consisting of reference images and modification texts. Although substantial progress has been made in recent years, existing methods assume that all samples are correctly matched. However, in real-world scenarios, due to high triplet annotation costs, CIR datasets inevitably contain annotation errors, resulting in incorrectly matched triplets. To address this issue, the problem of Noisy Triplet Correspondence (NTC) has attracted growing attention. We argue that noise in CIR can be categorized into two types: cross-modal correspondence noise and modality-inherent noise. The former arises from mismatches across modalities, whereas the latter originates from intra-modal background interference or visual factors irrelevant to the coarse-grained modification annotations. However, modality-inherent noise is often overlooked, and research on cross-modal correspondence noise remains nascent. To tackle above issues, we propose the Invariance and discrimiNaTion-awarE Noise neTwork (INTENT), comprising two components: Visual Invariant Composition and Bi-Objective Discriminative Learning, specifically designed to handle the two-aspect noise. The former applies causal intervention on the visual side via Fast Fourier Transform (FFT) to generate intervened composed features, enforcing visual invariance and enabling the model to ignore modality-inherent noise during composition. The latter adopts collaborative optimization with both positive and negative samples, and constructs a scalable decision boundary that dynamically adjusts decisions based on the loyalty degree, enabling robust correspondence discrimination. Extensive experiments on two widely used benchmark datasets demonstrate the superiority and robustness of INTENT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。