针对正负样本极不平衡的场景,提出更精准的正/未标记学习方法。
Focused PU learning from imbalanced data

- 用聚焦经验风险估计器,融合正例与未标记数据训练分类器。
- 在两种随机选正例设定下,性能超越现有方法,尤其对难区分正例有效。
- 适用于金融舞弊检测等真实场景,对小样本、难识别正例有显著提升。
我们提出一种在高度不平衡数据集上学习正例与未标记例(PU)的新方法。许多现实问题如疾病基因识别、定向营销、欺诈检测和推荐系统,因标注数据有限而难以用机器学习解决。训练数据通常包含正例和未标记实例,后者多为负例,但也包含若干正例。尽管PU学习已有研究,但很少方法能处理不平衡情况或识别与负例相似的难检测正例。我们的方法采用聚焦经验风险估计器,结合正例与未标记例来训练二分类器。实验表明,在两种标注机制——完全随机选择正例(SCAR)和随机选择正例(SAR)——下,该方法在不平衡数据集上达到领先性能。此外,我们在真实世界中的财务舞弊检测任务中验证了该方法的有效性。
原文摘要 · Abstract (English)
We propose a new method of learning from positive and unlabeled (PU) examples in highly imbalanced datasets. Many real-world problems, such as disease gene identification, targeted marketing, fraud detection, and recommender systems, are hard to address with machine learning methods, due to limited labeled data. Often, training data comprises positive and unlabeled instances, the latter typically being dominated by negative, but including also several positive instances. While PU learning is well-studied, few methods address imbalanced settings or hard-to-detect positive examples that resemble negative ones. Our approach uses a focused empirical risk estimator, incorporating both positive and unlabeled examples to train binary classifiers. Empirical evaluations demonstrate state-of-the-art performance on imbalanced datasets under two labeling mechanisms - selecting positives completely at random (SCAR) and selecting at random (SAR). Beyond these controlled experiments, we demonstrate the value of the proposed method in the real-world application of financial misstatement detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。