arXiv:2602.02081cs.LG2026-02

提出主动学习中正例与未标记数据的高效标注策略。

Active learning from positive and unlabeled examples

  • 通过自适应查询未标记样本,仅在正例且随机成功时获取标签。
  • 首次理论分析主动PU学习的标签复杂度,揭示标注效率下限。
  • 适合广告检测、异常识别等弱监督场景,兼顾成本与性能。

从正例与未标记数据中学习(PU学习)是一种弱监督二分类方法,其中学习者仅能获得部分正例的标签,其余样本均未标记。受广告投放和异常检测等应用启发,本文研究一种主动PU学习设置:学习者可自适应地从未标记数据池中查询实例,但只有当该实例为正例且独立的随机试验成功时,才会揭示其标签;否则学习者无法获得任何信息。本文首次对主动PU学习的标签复杂度进行了理论分析。

原文摘要 · Abstract (English)

Learning from positive and unlabeled data (PU learning) is a weakly supervised variant of binary classification in which the learner receives labels only for (some) positively labeled instances, while all other examples remain unlabeled. Motivated by applications such as advertising and anomaly detection, we study an active PU learning setting where the learner can adaptively query instances from an unlabeled pool, but a queried label is revealed only when the instance is positive and an independent coin flip succeeds; otherwise the learner receives no information. In this paper, we provide the first theoretical analysis of the label complexity of active PU learning.

主动学习弱监督标注效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。