arXiv:2502.02488cs.LG2025-02被引 2

现有图扩散模型生成的子结构分布与训练数据不符,新方法通过增强神经网络表达能力解决此问题。

Do Graph Diffusion Models Accurately Capture and Generate Substructure Distributions?

  • 以子结构频次为指标评估图扩散模型的分布拟合能力
  • 传统模型生成的新图子结构频次与训练集差异显著
  • 采用更强大图神经网络提升模型表达力,显著改善子结构生成质量

扩散模型在图生成任务中日益流行,但其对复杂图数据分布的学习能力尚不明确。不同于其他领域模型,主流图扩散模型如图变压器不具备普遍表达能力,难以准确建模图分布得分。本文聚焦子结构频次作为目标图分布的关键特征,发现现有模型在生成新图时无法保持训练集中观察到的子结构计数分布。为此,我们建立了图神经网络表达力与图扩散模型整体性能之间的理论联系,证明更强大的GNN骨干网络能更好捕捉复杂分布模式。通过将先进GNN集成至骨干架构,显著提升了子结构生成效果。

原文摘要 · Abstract (English)

Diffusion models have gained popularity in graph generation tasks; however, the extent of their expressivity concerning the graph distributions they can learn is not fully understood. Unlike models in other domains, popular backbones for graph diffusion models, such as Graph Transformers, do not possess universal expressivity to accurately model the distribution scores of complex graph data. Our work addresses this limitation by focusing on the frequency of specific substructures as a key characteristic of target graph distributions. When evaluating existing models using this metric, we find that they fail to maintain the distribution of substructure counts observed in the training set when generating new graphs. To address this issue, we establish a theoretical connection between the expressivity of Graph Neural Networks (GNNs) and the overall performance of graph diffusion models, demonstrating that more expressive GNN backbones can better capture complex distribution patterns. By integrating advanced GNNs into the backbone architecture, we achieve significant improvements in substructure generation.

图生成扩散模型子结构GNN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。