用GNN直接预测网络基序显著性,突破传统计数方法局限
Studying and Improving Graph Neural Network-based Motif Estimation
- 将基序显著性估计转为多目标回归,跳过子图计数
- 1-WL模型无法精确估计但能捕捉生成过程近似特征
- 为大规模图的基序分析提供可解释、稳定的新思路
图神经网络(GNN)是图表示学习的主要方法。然而,除了子图频率估计外,其在网络基序显著性-轮廓(SP)预测中的应用仍不充分,且文献中缺乏基准。我们提出将SP估计视为独立于子图频率估计的任务,从频率计数转向直接SP估计,并将其建模为多目标回归。该重构方法在可解释性、稳定性与大规模图上的可扩展性方面表现优异。我们在一个大规模合成数据集上验证方法,并进一步在真实网络上测试。实验表明,1-WL受限模型难以精确估计SP,但能通过对比预测SP与合成生成器输出的SP,近似捕捉网络生成过程。这是首次对基于GNN的基序估计的研究,暗示直接SP估计可突破子图计数在理论上的局限。
原文摘要 · Abstract (English)
Graph Neural Networks (GNNs) are a predominant method for graph representation learning. However, beyond subgraph frequency estimation, their application to network motif significance-profile (SP) prediction remains under-explored, with no established benchmarks in the literature. We propose to address this problem, framing SP estimation as a task independent of subgraph frequency estimation. Our approach shifts from frequency counting to direct SP estimation and modulates the problem as multitarget regression. The reformulation is optimised for interpretability, stability and scalability on large graphs. We validate our method using a large synthetic dataset and further test it on real-world graphs. Our experiments reveal that 1-WL limited models struggle to make precise estimations of SPs. However, they can generalise to approximate the graph generation processes of networks by comparing their predicted SP with the ones originating from synthetic generators. This first study on GNN-based motif estimation also hints at how using direct SP estimation can help go past the theoretical limitations that motif estimation faces when performed through subgraph counting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。