arXiv:2504.12651cs.LGcs.NE2025-04中稿 · GECCO 2025

针对少量正例的特征选择,利用数据聚类特性提升效果

Feature selection based on cluster assumption in PU learning

  • 基于聚类假设设计新方法,不把未标记数据当负例
  • 在合成数据和三个公开数据集上表现优于10种传统方法
  • 特别适合正例稀疏且呈聚集分布的现实场景

特征选择对高效数据挖掘至关重要,在正例-未标记(PU)学习场景中尤为关键,此时仅有少量正例标签,多数数据未标记。在某些真实世界任务中,经过良好特征选择的数据会形成正例集中分布的聚类。传统方法将未标记数据视为负例,难以捕捉正例的统计特性,导致性能不佳。为此,我们提出一种基于聚类假设的新型特征选择方法FSCPU,将特征选择建模为二元优化问题,目标函数显式融入了PU学习中的聚类假设。在合成数据上的实验验证了FSCPU在多种数据条件下的有效性。此外,在三个公开数据集上与10种传统算法对比,即使聚类假设不严格成立,FSCPU在下游分类任务中仍表现出竞争力。

原文摘要 · Abstract (English)

Feature selection is essential for efficient data mining and sometimes encounters the positive-unlabeled (PU) learning scenario, where only a few positive labels are available, while most data remains unlabeled. In certain real-world PU learning tasks, data subjected to adequate feature selection often form clusters with concentrated positive labels. Conventional feature selection methods that treat unlabeled data as negative may fail to capture the statistical characteristics of positive data in such scenarios, leading to suboptimal performance. To address this, we propose a novel feature selection method based on the cluster assumption in PU learning, called FSCPU. FSCPU formulates the feature selection problem as a binary optimization task, with an objective function explicitly designed to incorporate the cluster assumption in the PU learning setting. Experiments on synthetic datasets demonstrate the effectiveness of FSCPU across various data conditions. Moreover, comparisons with 10 conventional algorithms on three open datasets show that FSCPU achieves competitive performance in downstream classification tasks, even when the cluster assumption does not strictly hold.

特征选择PU学习聚类假设

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。