提出一种无需估计后验概率的类别先验估计方法,适用于正类样本不全的情况。
Prior shift estimation for positive unlabeled data through the lens of kernel embedding
- 基于核嵌入的分布匹配,直接求解先验估计问题。
- 在有限样本下表现稳定,优于或相当现有方法。
- 适合处理源数据中仅知正样本和总体分布的场景。
研究目标样本中类别先验的估计问题,其分布可能与源数据不同。假设源数据部分可观测:仅有正类样本和总体样本可用(即正类-无标签学习场景)。本文提出一种新型的直接先验估计器,避免了在两个群体中估计后验概率,具有简洁的几何解释。该方法基于再生核希尔伯特空间中的核嵌入与分布匹配技术,作为优化问题的显式解获得。我们建立了其渐近一致性,并给出了可实际计算的非渐近偏差上界。通过合成数据和真实数据的有限样本实验表明,该方法在性能上稳定优于或等同于现有方法。
原文摘要 · Abstract (English)
We study estimation of a class prior for unlabeled target samples which possibly differs from that of source population. Moreover, it is assumed that the source data is partially observable: only samples from the positive class and from the whole population are available (PU learning scenario). We introduce a novel direct estimator of a class prior which avoids estimation of posterior probabilities in both populations and has a simple geometric interpretation. It is based on a distribution matching technique together with kernel embedding in a Reproducing Kernel Hilbert Space and is obtained as an explicit solution to an optimisation task. We establish its asymptotic consistency as well as an explicit non-asymptotic bound on its deviation from the unknown prior, which is calculable in practice. We study finite sample behaviour for synthetic and real data and show that the proposal works consistently on par or better than its competitors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。