arXiv:2503.01361cond-mat.dis-nncs.LG2025-03被引 3

通过统计物理方法分析深度图卷积网络,揭示其接近最优分类性能的条件。

Statistical physics analysis of graph neural networks: Approaching optimality in the contextual stochastic block model

  • 用复制方法推导高维极限下的自由能,预测模型泛化性能。
  • 证明增大网络深度可逼近贝叶斯最优,且需相应调整架构防过度平滑。
  • 提出类动态平均场理论的约束方法,适用于深层图神经网络分析。

图神经网络(GNN)用于处理图结构数据,但其理论理解仍不充分。传统GCN在多步卷积中易因过度平滑导致远距离节点信息难以聚合。本文研究基于上下文随机块模型生成的数据上,基本GCN进行节点分类的泛化性能。在高维极限下,使用复制方法推导问题的自由能,预测其渐近性能。结果表明,增加卷积步数(即深度)对逼近贝叶斯最优至关重要;同时需随深度调整网络架构以避免过度平滑。最终的大深度极限趋近于贝叶斯最优,形成连续型GCN。技术上,采用类似动态平均场理论(DMFT)但施加初末时刻约束的方法处理连续极限,并通过大正则化展开求解性能方程。该方法为深层神经网络分析提供新工具。

原文摘要 · Abstract (English)

Graph neural networks (GNNs) are designed to process data associated with graphs. They are finding an increasing range of applications; however, as with other modern machine learning techniques, their theoretical understanding is limited. GNNs can encounter difficulties in gathering information from nodes that are far apart by iterated aggregation steps. This situation is partly caused by so-called oversmoothing; and overcoming it is one of the practically motivated challenges. We consider the situation where information is aggregated by multiple steps of convolution, leading to graph convolutional networks (GCNs). We analyze the generalization performance of a basic GCN, trained for node classification on data generated by the contextual stochastic block model. We predict its asymptotic performance by deriving the free energy of the problem, using the replica method, in the high-dimensional limit. Calling depth the number of convolutional steps, we show the importance of going to large depth to approach the Bayes-optimality. We detail how the architecture of the GCN has to scale with the depth to avoid oversmoothing. The resulting large depth limit can be close to the Bayes-optimality and leads to a continuous GCN. Technically, we tackle this continuous limit via an approach that resembles dynamical mean-field theory (DMFT) with constraints at the initial and final times. An expansion around large regularization allows us to solve the corresponding equations for the performance of the deep GCN. This promising tool may contribute to the analysis of further deep neural networks.

图神经网络统计物理深度学习泛化性能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。