arXiv:2506.00244cs.LGstat.ML2025-06被引 2

用留一法影响函数识别图数据噪声标签,提升节点分类鲁棒性

DeGLIF for Label Noise Robust Node Classification using GNNs

  • 基于留一法影响函数估计删去节点对验证损失的影响
  • 无需噪声模型信息,在多个数据集上准确率超越基线方法
  • 提出两种不依赖噪声先验的检测变体,理论证明可提升风险

噪声标签数据相比清洁标签数据成本更低,图数据亦是如此。本文提出一种去噪技术DeGLIF:利用留一法影响函数进行图数据去噪。DeGLIF结合少量干净数据与留一法影响函数,实现图数据上对标签噪声鲁棒的节点级预测。留一法影响函数用于近似移除训练样本后模型参数的变化,近期研究已实现对图神经网络(GNNs)的留一法影响函数计算。本文进一步扩展该方法,以估计移除训练节点后验证损失的变化。利用此估计值及一个新提出的理论驱动重标注函数,完成训练数据集的去噪。本文提出两种DeGLIF变体来识别噪声节点,二者均无需噪声模型或噪声水平信息,且不估计这些量。其中一个变体的噪声点检测结果被证明可实际增加风险。在多种数据集上进行了详尽的计算实验,验证了DeGLIF的有效性,其准确率优于其他基线算法。

原文摘要 · Abstract (English)

Noisy labelled datasets are generally inexpensive compared to clean labelled datasets, and the same is true for graph data. In this paper, we propose a denoising technique DeGLIF: Denoising Graph Data using Leave-One-Out Influence Function. DeGLIF uses a small set of clean data and the leave-one-out influence function to make label noise robust node-level prediction on graph data. Leave-one-out influence function approximates the change in the model parameters if a training point is removed from the training dataset. Recent advances propose a way to calculate the leave-one-out influence function for Graph Neural Networks (GNNs). We extend that recent work to estimate the change in validation loss, if a training node is removed from the training dataset. We use this estimate and a new theoretically motivated relabelling function to denoise the training dataset. We propose two DeGLIF variants to identify noisy nodes. Both these variants do not require any information about the noise model or the noise level in the dataset; DeGLIF also does not estimate these quantities. For one of these variants, we prove that the noisy points detected can indeed increase risk. We carry out detailed computational experiments on different datasets to show the effectiveness of DeGLIF. It achieves better accuracy than other baseline algorithms

图神经网络标签噪声去噪鲁棒学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。