arXiv:2501.10935cs.CVcs.AI2025-01AAAI被引 5

提出三模型协作框架,提升图像文本检索在噪声数据下的鲁棒性。

TSVC:Tripartite Learning with Semantic Variation Consistency for Robust Image-Text Retrieval

  • 设计协调器、主模型与助手模型的三元协同机制,增强数据视角多样性。
  • 在噪声率升至30%时,仍保持较高检索准确率,优于现有方法。
  • 适用于标注不精准的跨模态检索场景,尤其适合数据质量差的项目。

跨模态检索通过语义相关性映射不同模态的数据。现有方法隐式假设数据对齐良好,忽视普遍存在的标注噪声(即噪声对应关系,NC),导致性能下降。尽管已有研究采用同构结构的共教范式提供不同数据视角,但模型差异主要源于随机初始化,训练过程使模型趋于同质化,附加信息有限。为此,本文提出三元协同学习框架TSVC,包含协调器、主模型与助手模型。协调器分发数据,助手模型为带噪声标签的样本提供多样化支持。引入基于互信息变化的软标签估计方法,量化新样本中的噪声并分配对应软标签。同时设计新损失函数,提升鲁棒性与训练效率。在三个常用数据集上的大量实验表明,即使噪声比例增至30%,TSVC仍显著优于基线,检索精度更高且训练稳定。

原文摘要 · Abstract (English)

Cross-modal retrieval maps data under different modality via semantic relevance. Existing approaches implicitly assume that data pairs are well-aligned and ignore the widely existing annotation noise, i.e., noisy correspondence (NC). Consequently, it inevitably causes performance degradation. Despite attempts that employ the co-teaching paradigm with identical architectures to provide distinct data perspectives, the differences between these architectures are primarily stemmed from random initialization. Thus, the model becomes increasingly homogeneous along with the training process. Consequently, the additional information brought by this paradigm is severely limited. In order to resolve this problem, we introduce a Tripartite learning with Semantic Variation Consistency (TSVC) for robust image-text retrieval. We design a tripartite cooperative learning mechanism comprising a Coordinator, a Master, and an Assistant model. The Coordinator distributes data, and the Assistant model supports the Master model's noisy label prediction with diverse data. Moreover, we introduce a soft label estimation method based on mutual information variation, which quantifies the noise in new samples and assigns corresponding soft labels. We also present a new loss function to enhance robustness and optimize training effectiveness. Extensive experiments on three widely used datasets demonstrate that, even at increasing noise ratios, TSVC exhibits significant advantages in retrieval accuracy and maintains stable training performance.

跨模态检索噪声鲁棒三元学习软标签

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。