针对图数据类别不平衡问题,提出不确定性感知伪标签方法提升分类准确率。
UPL: Uncertainty-aware Pseudo-labeling for Imbalance Transductive Node Classification
- 基于不确定性感知机制优化伪标签分配,降低噪声影响。
- 在多个基准数据集上显著优于现有最优方法,提升分类性能。
- 适合处理标签分布不均的图节点分类任务,尤其对少数类效果明显。
图结构数据常面临类别不平衡问题,使节点分类任务更加复杂。本文首先为不平衡的直推式节点分类提供了总体风险的上界分析,随后提出一种简单而新颖的算法——不确定性感知伪标签(UPL)。该方法通过为未标记节点分配伪标签来缓解不平衡带来的负面影响。此外,UPL通过一种新颖的不确定性感知策略,减少伪标签训练过程中的噪声,从而提升伪标签准确性。我们在多个基准数据集上全面评估了UPL算法,结果表明其性能显著优于现有最先进方法。
原文摘要 · Abstract (English)
Graph-structured datasets often suffer from class imbalance, which complicates node classification tasks. In this work, we address this issue by first providing an upper bound on population risk for imbalanced transductive node classification. We then propose a simple and novel algorithm, Uncertainty-aware Pseudo-labeling (UPL). Our approach leverages pseudo-labels assigned to unlabeled nodes to mitigate the adverse effects of imbalance on classification accuracy. Furthermore, the UPL algorithm enhances the accuracy of pseudo-labeling by reducing training noise of pseudo-labels through a novel uncertainty-aware approach. We comprehensively evaluate the UPL algorithm across various benchmark datasets, demonstrating its superior performance compared to existing state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。