arXiv:2606.11689cs.CV2026-06中稿 · ICMR 2026

通过结构感知与动态值校准,提升图像检索在噪声数据下的鲁棒性。

RankVR: Low-Rank Structure Perception and Value Recalibration for Robust Composed Image Retrieval

论文配图:RankVR: Low-Rank Structure Perception and Value Recalibration for Robust Composed Image Retrieval
图 1 · 摘自论文原文
  • 利用相关矩阵的秩差异识别破坏语义对称性的噪声样本。
  • 动态量化三元组语义价值,在保留困难正例的同时抑制逻辑冲突噪声。
  • 适用于存在大量噪声标注的复杂图像检索任务,尤其适合工业级应用。

组合图像检索(CIR)要求模型对参考图像和修改文本进行联合推理。然而大规模数据集中的噪声三元组对应(NTC)严重制约模型性能。现有去噪方法或仅处理二元错配,或依赖标量点估计,忽视了样本群体间的全局结构相关性及训练过程中的动态值变化,导致效果不佳。本文识别出两大未解难题:语义关联的全局结构不一致、困难样本判别不确定性。为此提出RankVR框架,通过全局结构一致性与动态值感知构建稳健的CIR模型。首先设计全局结构一致性感知(GSCP)模块,利用相关矩阵的有效秩分离干净样本与结构噪声,通过秩差识别破坏宏观语义对称性的样本;其次开发自适应语义值校准(ASVC)模块,结合训练潜力与可靠性动态量化每个三元组的语义价值,确保困难正例有效利用,同时抑制逻辑冲突噪声。在FashionIQ和CIRR基准上的大量实验表明,RankVR显著优于现有最先进方法,验证其在噪声环境下的卓越鲁棒性。

原文摘要 · Abstract (English)

Composed Image Retrieval (CIR) constitutes a pivotal paradigm requiring models to perform joint reasoning on reference images and modification texts. However, the prevalence of Noisy Triplet Correspondence (NTC) in large-scale datasets severely constrains model performance. Existing denoising methods either target binary mismatches or rely on scalar-based point-wise estimation, neglecting rich global structural correlations among sample populations and dynamic value variations during training, thereby yielding suboptimal results. This paper identifies two critical unresolved challenges: Global Structural Inconsistency of Semantic Correlations and Hard Sample Discrimination Uncertainty. To address these, we propose RankVR, a framework designed to construct a robust CIR model via global structure consistency and dynamic value perception. Specifically, we introduce the Global Structure Consistency Perception (GSCP) module, which utilizes the Effective Rank of the Correlation Matrix to decouple clean samples from structural noise. By measuring rank difference, GSCP identifies samples disrupting macroscopic semantic symmetry. Furthermore, we develop the Adaptive Semantic Value Calibration (ASVC) module to distinguish high-value hard clean samples. By integrating training potential and reliability, it dynamically quantifies the semantic value of each triplet, ensuring effective utilization of hard samples while suppressing noise characterized by logical conflicts. Extensive experiments on the FashionIQ and CIRR benchmark datasets demonstrate that RankVR significantly outperforms existing state-of-the-art methods, validating its superior robustness in noisy environments.

图像检索去噪结构感知鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。