不依赖先验正例比例,给出正例-无标签学习的样本量理论边界。
Learning from positive and unlabeled examples -Finite size sample bounds
- 在不假设正例比例已知的前提下分析PU学习的统计复杂性
- 推导出正例和无标签数据所需样本量的上下界
- 为实际应用提供理论指导,适合关注理论保障的研究者
正例-无标签(PU)学习是监督分类的一种变体,其中学习者仅能获取正例样本的标签。该问题在诸多真实场景中出现。现有方法通常依赖两个简化假设:正例训练数据来自正例分布的限制,或正例比例(即类别先验)已知。本文对更广泛设置下的PU学习统计复杂性进行了理论分析。不同于以往工作,本研究未假设学习者知晓类别先验。我们证明了正例与无标签样本所需样本量的上下界。
原文摘要 · Abstract (English)
PU (Positive Unlabeled) learning is a variant of supervised classification learning in which the only labels revealed to the learner are of positively labeled instances. PU learning arises in many real-world applications. Most existing work relies on the simplifying assumptions that the positively labeled training data is drawn from the restriction of the data generating distribution to positively labeled instances and/or that the proportion of positively labeled points (a.k.a. the class prior) is known apriori to the learner. This paper provides a theoretical analysis of the statistical complexity of PU learning under a wider range of setups. Unlike most prior work, our study does not assume that the class prior is known to the learner. We prove upper and lower bounds on the required sample sizes (of both the positively labeled and the unlabeled samples).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。