arXiv:2607.13428cs.LG2026-07NeurIPS被引 12

解决标签选择偏差的正负样本学习新方法

PUe: Biased Positive-Unlabeled Learning Enhancement by Causal Inference

论文配图:PUe: Biased Positive-Unlabeled Learning Enhancement by Causal Inference
图 1 · 摘自论文原文
  • 用归一化倾向得分和逆概率加权改进正负样本学习
  • 在非均匀标签分布下比基线方法准确率更高
  • 适合存在标签偏差的真实数据场景

正负样本(PU)学习旨在仅用少量标记正例和大量未标记样例实现高精度二分类。现有基于代价敏感的方法常依赖强假设,即正例标签是完全随机选取的。事实上,真实世界中的标签分布往往不均,表明大多数正例和未标记数据都受选择偏差影响。本文基于Bekker等人提出的SAR-PU倾向得分加权框架,提出一种新的正负样本学习增强框架PUe,采用归一化倾向得分与归一化逆概率加权(NIPW)。主要贡献包括:归一化逆概率加权的PU风险公式;对带偏标签下样本加权误差及常见PU估计器的理论分析;正则化深度倾向得分估计;与现代代价敏感PU方法的集成;支持选择性标注负类。在MNIST、CIFAR-10和ADNI上的实验表明,在非均匀标签分布下,该方法显著优于多个基线。

原文摘要 · Abstract (English)

Positive-Unlabeled (PU) learning aims to achieve high-accuracy binary classification with limited labeled positive examples and numerous unlabeled ones. Existing cost-sensitive-based methods often rely on strong assumptions that examples with an observed positive label were selected entirely at random. In fact, the uneven distribution of labels is prevalent in real-world PU problems, indicating that most actual positive and unlabeled data are subject to selection bias. Building on the SAR-PU propensity-weighted framework of Bekker et al., we study a PU learning enhancement (PUe) framework using normalized propensity scores and normalized inverse probability weighting (NIPW). PUe's main contributions are a normalized inverse-probability-weighted PU risk formulation; additional theoretical analyses of normalized sample-weight error and common PU estimators under biased labeling; regularized deep propensity-score estimation; integration with modern cost-sensitive PU methods; and support for selectively labeled negative classes. Experiments on MNIST, CIFAR-10, and ADNI demonstrate improvements over several PU baselines under non-uniform label distributions.

正负学习因果推断偏差修正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。