arXiv:2604.19386cs.CV2026-04中稿 · CVPR被引 16

解决图像检索中模糊匹配导致的模型污染问题,提升多模态查询准确性。

Air-Know: Arbiter-Calibrated Knowledge-Internalizing Robust Network for Composed Image Retrieval

论文配图:Air-Know: Arbiter-Calibrated Knowledge-Internalizing Robust Network for Composed Image Retrieval
图 1 · 摘自论文原文
  • 用大模型构建精准锚点数据集,分离专家与判别器角色。
  • 在噪声数据下比当前最佳方法提升12.3%准确率,传统任务也表现优异。
  • 适合做多模态图像检索且关注鲁棒性的研究者和开发者。

组合图像检索(CIR)因其灵活的多模态查询方式受到广泛关注,但其发展严重受限于噪声三元组对应(NTC)问题。现有鲁棒学习方法依赖“小损失假设”,但NTC特有的语义模糊性(如部分匹配)破坏了该假设,导致噪声识别不可靠,使模型陷入学习者与判别器相互依赖的恶性循环,最终引发灾难性“表征污染”。为此,我们提出新型“专家-代理-分流”解耦范式Air-Know(ArbIteR calibrated Knowledge iNternalizing rObust netWork)。Air-Know包含三个核心模块:(1) 外部先验仲裁(EPA),利用多模态大语言模型(MLLMs)作为离线专家构建高精度锚点数据集;(2) 专家知识内化(EKI),高效引导轻量级代理“判别器”内化专家的判别逻辑;(3) 双流协调(DSR),利用EKI的匹配置信度分流训练数据,实现纯净对齐流与表征反馈协调流。在多个CIR基准数据集上的实验表明,Air-Know在NTC设置下显著优于现有最先进方法,同时在传统CIR任务中也展现出强大竞争力。

原文摘要 · Abstract (English)

Composed Image Retrieval (CIR) has attracted significant attention due to its flexible multimodal query method, yet its development is severely constrained by the Noisy Triplet Correspondence (NTC) problem. Most existing robust learning methods rely on the "small loss hypothesis", but the unique semantic ambiguity in NTC, such as "partial matching", invalidates this assumption, leading to unreliable noise identification. This entraps the model in a self dependent vicious cycle where the learner is intertwined with the arbiter, ultimately causing catastrophic "representation pollution". To address this critical challenge, we propose a novel "Expert-Proxy-Diversion" decoupling paradigm, named Air-Know (ArbIteR calibrated Knowledge iNternalizing rObust netWork). Air-Know incorporates three core modules: (1) External Prior Arbitration (EPA), which utilizes Multimodal Large Language Models (MLLMs) as an offline expert to construct a high precision anchor dataset; (2) Expert Knowledge Internalization (EKI), which efficiently guides a lightweight proxy "arbiter" to internalize the expert's discriminative logic; (3) Dual Stream Reconciliation (DSR), which leverages the EKI's matching confidence to divert the training data, achieving a clean alignment stream and a representation feedback reconciliation stream. Extensive experiments on multiple CIR benchmark datasets demonstrate that Air-Know significantly outperforms existing SOTA methods under the NTC setting, while also showing strong competitiveness in traditional CIR.

图像检索多模态鲁棒学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。