arXiv:2509.22868cs.LGstat.ML2025-09

不同邻居采样方式训练出的GNN模型本质不同,影响预测性能。

Neighborhood Sampling Does Not Learn the Same Graph Neural Network

  • 用神经正切核分析多种邻居采样方法的理论行为
  • 小样本下各方法后验分布差异大,预测误差不可比
  • 解释为何无一种采样策略在所有场景占优

邻居采样是大规模图神经网络训练中的关键环节,能抑制邻域规模随层数指数增长,保持可接受的内存与时间开销。尽管已成为实际标准,其系统性行为仍不清晰。本文利用神经正切核工具,研究几种典型邻居采样方法及其对应的后验高斯过程。在样本量有限时,不同采样方法得到的后验分布各不相同,虽随样本量增加趋于一致,但后验协方差无法比较,这与观察到的无任何一种采样方法始终占优的现象一致。

原文摘要 · Abstract (English)

Neighborhood sampling is an important ingredient in the training of large-scale graph neural networks. It suppresses the exponential growth of the neighborhood size across network layers and maintains feasible memory consumption and time costs. While it becomes a standard implementation in practice, its systemic behaviors are less understood. We conduct a theoretical analysis by using the tool of neural tangent kernels, which characterize the (analogous) training dynamics of neural networks based on their infinitely wide counterparts -- Gaussian processes (GPs). We study several established neighborhood sampling approaches and the corresponding posterior GP. With limited samples, the posteriors are all different, although they converge to the same one as the sample size increases. Moreover, the posterior covariance, which lower-bounds the mean squared prediction error, is uncomparable, aligning with observations that no sampling approach dominates.

图神经网络邻居采样理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。