arXiv:2409.07712cs.SIcs.LG2024-09被引 1

通过生成新节点提升稀疏标签图的分类性能

Virtual Node Generation for Node Classification in Sparsely-Labeled Graphs

  • 设计优化问题生成高质量合成节点,增强标签传播
  • 在10个数据集上显著优于14种基线方法
  • 兼容主流图学习方法,适合标签稀缺场景

在机器学习领域,数据生成方法通过扩充稀疏标签来生成更多有信息量的训练样本,表现优异。但在图结构中,由于节点间复杂的依赖关系,此类方法研究较少。本文提出一种新型节点生成方法,将少量高质量合成节点注入图中作为额外标签节点,以最优方式扩展标签信息的传播。该框架不依赖特定图学习或下游分类技术,可与多数主流图预训练(自监督学习)、半监督学习及元学习方法兼容。核心贡献在于通过求解一个新颖的优化问题设计生成节点的位置:(1) 最小化分类损失,保证训练准确率;(2) 最大化向低置信度节点传播标签,确保高质量传播。理论上证明,这种双重优化能最大化节点分类的全局置信度。实验表明,在10个公开数据集上,性能显著优于14种基线方法。

原文摘要 · Abstract (English)

In the broader machine learning literature, data-generation methods demonstrate promising results by generating additional informative training examples via augmenting sparse labels. Such methods are less studied in graphs due to the intricate dependencies among nodes in complex topology structures. This paper presents a novel node generation method that infuses a small set of high-quality synthesized nodes into the graph as additional labeled nodes to optimally expand the propagation of labeled information. By simply infusing additional nodes, the framework is orthogonal to the graph learning and downstream classification techniques, and thus is compatible with most popular graph pre-training (self-supervised learning), semi-supervised learning, and meta-learning methods. The contribution lies in designing the generated node set by solving a novel optimization problem. The optimization places the generated nodes in a manner that: (1) minimizes the classification loss to guarantee training accuracy and (2) maximizes label propagation to low-confidence nodes in the downstream task to ensure high-quality propagation. Theoretically, we show that the above dual optimization maximizes the global confidence of node classification. Our Experiments demonstrate statistically significant performance improvements over 14 baselines on 10 publicly available datasets.

图神经网络数据生成标签传播

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。