首个系统化评估正例-未标记学习算法的基准框架
Accessible, Realistic, and Fair Evaluation of Positive-Unlabeled Learning Algorithms
- 构建首个统一的PU学习评估基准,解决实验设置不一致问题
- 发现并修正一样本场景中的标签偏移问题,提升评估公平性
- 适合算法研究者和需要可靠评估的工业应用开发者
正例-未标记(PU)学习是一种弱监督二分类问题,目标是从仅有的正例和未标记数据中训练分类器,无法获取负例数据。近年来虽涌现大量PU学习算法,但实验设置差异大,难以判断性能优劣。本文提出首个系统性评估PU学习算法的基准框架。在实现过程中,我们识别出影响评估真实性和公平性的关键因素:一方面,许多算法依赖包含负例的验证集进行模型选择,这在传统PU设定下不现实;我们系统研究了适用于PU学习的模型选择标准。另一方面,PU学习存在一样本与两样本两种设定,现有评估协议严重偏向一样本设置,并忽视两者差异。我们揭示了一样本设置中未标记数据内部的标签偏移问题,提出一种简单有效的校准方法,确保同族及跨族间的公平比较。本框架旨在为未来PU学习算法提供可访问、真实且公平的评估环境。
原文摘要 · Abstract (English)
Positive-unlabeled (PU) learning is a weakly supervised binary classification problem, in which the goal is to learn a binary classifier from only positive and unlabeled data, without access to negative data. In recent years, many PU learning algorithms have been developed to improve model performance. However, experimental settings are highly inconsistent, making it difficult to identify which algorithm performs better. In this paper, we propose the first PU learning benchmark to systematically compare PU learning algorithms. During our implementation, we identify subtle yet critical factors that affect the realistic and fair evaluation of PU learning algorithms. On the one hand, many PU learning algorithms rely on a validation set that includes negative data for model selection. This is unrealistic in traditional PU learning settings, where no negative data are available. To handle this problem, we systematically investigate model selection criteria for PU learning. On the other hand, PU learning involves different problem settings and corresponding solution families, i.e., the one-sample and two-sample settings. However, existing evaluation protocols are heavily biased towards the one-sample setting and neglect the significant difference between them. We identify the internal label shift problem of unlabeled training data for the one-sample setting and propose a simple yet effective calibration approach to ensure fair comparisons within and across families. We hope our framework will provide an accessible, realistic, and fair environment for evaluating PU learning algorithms in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。