arXiv:2410.13439cs.LGcs.CL2024-10中稿 · Transactions on Ma…被引 2

提出新损失函数,让多标签对比学习更准更稳。

Similarity-Dissimilarity Loss for Multi-label Supervised Contrastive Learning

  • 基于相似与相异关系动态重加权样本,优化正负样本定义。
  • 在图像、文本及医疗数据上均超越现有方法,提升显著。
  • 理论证明有效,兼容单标签与多标签场景。

监督对比学习通过利用标签信息取得了显著成果;然而,在多标签场景中确定正样本仍是关键挑战。在多标签监督对比学习(MSCL)中,多标签关系尚未明确定义,导致正样本识别模糊以及对比损失函数构建困难。为此,本文:(i) 系统化地定义了MSCL中的多标签关系;(ii) 提出一种新颖的相似-相异损失(Similarity-Dissimilarity Loss),根据样本间的相似性和相异性动态重加权;(iii) 通过严谨的数学分析提供理论支持,验证方法的合理性与有效性;(iv) 统一了单标签与多标签监督对比损失的形式与范式。我们在图像、文本模态及医学领域进行了实验,结果表明,该方法在综合评估中持续优于基线模型,展现出更强的有效性与鲁棒性。

原文摘要 · Abstract (English)

Supervised contrastive learning has achieved remarkable success by leveraging label information; however, determining positive samples in multi-label scenarios remains a critical challenge. In multi-label supervised contrastive learning (MSCL), multi-label relations are not yet fully defined, leading to ambiguity in identifying positive samples and formulating contrastive loss functions to construct the representation space. To address these challenges, we: (i) systematically formulate multi-label relations in MSCL, (ii) propose a novel Similarity-Dissimilarity Loss, which dynamically re-weights samples based on similarity and dissimilarity factors, (iii) further provide theoretically grounded proofs for our method through rigorous mathematical analysis that supports the formulation and effectiveness, and (iv) offer a unified form and paradigm for both single-label and multi-label supervised contrastive loss. We conduct experiments on both image and text modalities and further extend the evaluation to the medical domain. The results show that our method consistently outperforms baselines in comprehensive evaluations, demonstrating its effectiveness and robustness.

对比学习多标签损失函数图像识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。