arXiv:2510.25226cs.LGcs.AI2025-10

解决多分类正例-未标记学习中的偏差风险估计问题

Cost-Sensitive Unbiased Risk Estimation for Multi-Class Positive-Unlabeled Learning

  • 用自适应损失加权实现无偏风险估计
  • 在8个数据集上准确率与稳定性均优于基线
  • 适合负样本难获取的多分类场景

正例-未标记(PU)学习面对只有正例和未标记数据的情况,负例缺失或未标注。这在实际应用中常见,因标注可靠负例困难或成本高。尽管PU学习进展显著,多分类情况(MPU)仍具挑战:许多方法无法保证无偏风险估计,影响性能与稳定性。本文提出一种基于自适应损失加权的成本敏感多分类PU方法。在经验风险最小化框架下,为正例与推断负例(从未标记混合中得出)的损失分配不同、数据依赖的权重,使目标风险的无偏估计成为可能。我们形式化了MPU数据生成过程,并建立了所提估计器的一般化误差界。在八个公开数据集上的实验表明,覆盖不同类别先验和类别数,结果在准确率与稳定性上均持续优于强基线。

原文摘要 · Abstract (English)

Positive--Unlabeled (PU) learning considers settings in which only positive and unlabeled data are available, while negatives are missing or left unlabeled. This situation is common in real applications where annotating reliable negatives is difficult or costly. Despite substantial progress in PU learning, the multi-class case (MPU) remains challenging: many existing approaches do not ensure \emph{unbiased risk estimation}, which limits performance and stability. We propose a cost-sensitive multi-class PU method based on \emph{adaptive loss weighting}. Within the empirical risk minimization framework, we assign distinct, data-dependent weights to the positive and \emph{inferred-negative} (from the unlabeled mixture) loss components so that the resulting empirical objective is an unbiased estimator of the target risk. We formalize the MPU data-generating process and establish a generalization error bound for the proposed estimator. Extensive experiments on \textbf{eight} public datasets, spanning varying class priors and numbers of classes, show consistent gains over strong baselines in both accuracy and stability.

PU学习多分类无偏估计自适应加权

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。