arXiv:2412.02240cs.LG2024-12被引 2

提出ESA方法,解决多正例无标签学习中的风险偏移问题。

ESA: Example Sieve Approach for Multi-Positive and Unlabeled Learning

  • 通过示例筛选机制,利用每个样本的确定性损失值剔除低质量样本。
  • 理论证明其风险估计误差达到最优参数收敛速率。
  • 在多个真实数据集上优于现有方法,适合复杂标注场景。

从多正例和无标签(MPU)数据中学习逐渐受到实际应用关注。然而,当模型灵活性较高时,MPU面临最小风险偏移的问题,如图 ef{moti}所示。为缓解该问题,本文提出示例筛法(ESA),通过训练阶段各样本的确定性损失(CL)值筛选训练样本,以构建多分类器。我们分析了所提风险估计器的一致性,并证明其估计误差达到最优参数收敛速率。大量实验证明,该方法在多个真实数据集上均优于先前方法。

原文摘要 · Abstract (English)

Learning from Multi-Positive and Unlabeled (MPU) data has gradually attracted significant attention from practical applications. Unfortunately, the risk of MPU also suffer from the shift of minimum risk, particularly when the models are very flexible as shown in Fig.\ref{moti}. In this paper, to alleviate the shifting of minimum risk problem, we propose an Example Sieve Approach (ESA) to select examples for training a multi-class classifier. Specifically, we sieve out some examples by utilizing the Certain Loss (CL) value of each example in the training stage and analyze the consistency of the proposed risk estimator. Besides, we show that the estimation error of proposed ESA obtains the optimal parametric convergence rate. Extensive experiments on various real-world datasets show the proposed approach outperforms previous methods.

多正例学习无标签学习风险偏移示例筛选

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。