小图训练的神经微分图模型可直接部署到更大相似图上,理论证明其有效性。
Zero-Shot Size Transfer for Neural ODEs on Sparse Random Graphs: Graphon Limits and Adjoint Convergence

- 基于图随机极限理论,建立神经微分图模型在稀疏图上的零样本扩展框架
- 在稀疏随机图上,解轨迹收敛速度达O((α_n n)^{-1/2}),高概率成立
- 实验证明跨图类零样本迁移有效,适用于图规模变化场景
图神经微分方程(GNDE)通过图神经网络参数化神经微分方程的速度场,建模连续时间图动态。其局部、与规模无关的滤波器特性暗示了零样本规模迁移原则:在小图上训练后可直接部署于更大、相似的图上而无需重训练。本文针对从图随机过程(graphons)采样的稀疏随机图,建立了该原则的定量理论。考虑图随机神经微分方程(Graphon-NDE)及其伴随系统(adjoint Graphon-NDE),作为前向与伴随GNDE系统的无限节点极限,并证明其适定性。对于包含n个节点、稀疏参数为α_n的随机图,我们证明了GNDE解以概率接近1的方式,以速率O((α_n n)^{-1/2})收敛至Graphon-NDE解,忽略对数因子。同时建立了伴随系统(控制隐状态与参数梯度)的全局时间一致收敛界。进一步研究离散化-再优化(DTO)与优化-再离散化(OTD)训练方法。在显式欧拉离散化下,使用M步时,证明了两种方法渐近一致,隐状态与局部参数梯度差异分别达到O(1/M)与O(1/M^2),忽略稀疏性与对数因子影响。在HSBM和帐篷图随机过程上的实验支持理论速率,零样本迁移实验在四个图随机过程类别上展示了学习后的GNDE在独立采样大图上的准确部署能力。
原文摘要 · Abstract (English)
Graph Neural Differential Equations (GNDEs) model continuous-time graph dynamics by parameterizing Neural ODE velocity fields with Graph Neural Networks. Their local, size-independent filters suggest a zero-shot size-transfer principle: train on a small graph and deploy on larger, similar graphs without retraining. We develop a quantitative theory for this principle on sparse random graphs sampled from graphons. We consider Graphon Neural Differential Equations (Graphon-NDEs) and adjoint Graphon-NDEs as the infinite-node limits of the forward and adjoint GNDE systems, and establish well-posedness. For an $n$-node random graph with sparsity parameter $α_n$, we prove trajectory-wise convergence of GNDE solutions to Graphon-NDE solutions at rate $O((α_n n)^{-1/2})$, up to logarithmic factors, with high probability. We also establish uniform-in-time convergence bounds for adjoint systems governing hidden-state and parameter gradients. We further study discretize-then-optimize (DTO) and optimize-then-discretize (OTD) training. Under explicit Euler discretization with $M$ steps, we show that DTO and OTD are asymptotically consistent, with hidden-state and local parameter-gradient discrepancies of orders $O(1/M)$ and $O(1/M^2)$, respectively, up to sparsity and logarithmic factors. Experiments on HSBM and tent graphons support the theoretical rates, while zero-shot transfer experiments across four graphon classes demonstrate accurate deployment of learned GNDEs on larger independently sampled graphs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。