提出密度感知方法,让半监督学习更好利用数据分布结构。
Probability-density-aware Semi-supervised Learning
- 引入概率密度感知度量,更准确识别邻近点相似性
- 新算法PMLP在多个数据集上超越主流方法,最高提升3.2%准确率
- 理论证明伪标签是该方法的特例,统一了现有框架
半监督学习(SSL)通常依赖邻近点同类别(邻近假设)和不同簇点属不同类(聚类假设)的先验。现有方法多通过相似度度量寻找邻近点,忽视聚类假设,导致未充分利用无标签数据。本文首次系统研究概率密度在SSL中的关键作用,为聚类假设奠定理论基础。为此,提出概率密度感知度量(PM),用于区分邻近点间的相似性。为进一步改进标签传播,设计概率密度感知度量标签传播(PMLP)算法,充分考虑聚类假设。最后,证明传统伪标签可视为PMLP的特例,提供对PMLP优越性能的完整理论解释。大量实验表明,PMLP在多个基准数据集上表现优异,显著优于现有方法。
原文摘要 · Abstract (English)
Semi-supervised learning (SSL) assumes that neighbor points lie in the same category (neighbor assumption), and points in different clusters belong to various categories (cluster assumption). Existing methods usually rely on similarity measures to retrieve the similar neighbor points, ignoring cluster assumption, which may not utilize unlabeled information sufficiently and effectively. This paper first provides a systematical investigation into the significant role of probability density in SSL and lays a solid theoretical foundation for cluster assumption. To this end, we introduce a Probability-Density-Aware Measure (PM) to discern the similarity between neighbor points. To further improve Label Propagation, we also design a Probability-Density-Aware Measure Label Propagation (PMLP) algorithm to fully consider the cluster assumption in label propagation. Last but not least, we prove that traditional pseudo-labeling could be viewed as a particular case of PMLP, which provides a comprehensive theoretical understanding of PMLP's superior performance. Extensive experiments demonstrate that PMLP achieves outstanding performance compared with other recent methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。