提出可扩展的时序图聚类方法,融合谱聚类与社区检测理论,提升效率与精度。
Learning and Clustering on Temporal Graphs: Principles, Primitives, and Pooling
- 基于谱理论与随机块模型,统一图学习与社区检测框架
- 利用GPU加速实现多层时序聚类,处理大规模数据更高效
- 适用于无属性或弱属性场景,指导神经模型与传统算法的选择
本文聚焦时序图上的学习问题,尤其关注聚类任务:通过聚合节点、边及时间动态信息,获得粗粒度表示,该任务与图神经网络中的池化操作或网络科学中的社区检测相关。尽管图神经网络在多数下游任务中表现优异,但其相对于经典描述性与推断性聚类算法的优势仍不明确,尤其是在效率与恢复准确率方面。本文从三个关联视角展开:原理层面,通过共享谱基础与随机块模型中的可检测性阈值,连接图学习与社区检测;原语层面,借助GPU加速的时序后端,使谱聚类与多层模块度优化可计算;池化层面,将有理论依据的社区检测视为时序图上的粗粒化算子。结果表明,在缺乏或属性较弱时,传统算法仍更适用——瓶颈在于可扩展性而非准确性;而当结构、时间与属性信号一致时,神经模型更具优势。通过实现可扩展的时序聚类,GPU加速的原语为理论驱动的池化提供了路径,同时提出核心问题:基于社区的粗粒化能否保留下游学习所需的动态特性?
原文摘要 · Abstract (English)
This work focuses on the problem of learning on temporal graphs, with particular emphasis on the task of clustering: obtaining coarse-grained representations by aggregating information from nodes, edges, and temporal dynamics - a task related to pooling in machine learning on graphs, or community detection in network science. Although graph neural networks reach state-of-the-art performance across many downstream graph tasks, their advantage over established descriptive and inferential clustering algorithms is far less settled, especially under demands of efficiency and recovery accuracy. We frame this tension through three linked perspectives: principles, connecting graph learning and community detection through shared spectral foundations and detectability thresholds in stochastic block model regimes; primitives, making spectral clustering and multislice modularity optimization tractable through GPU-accelerated temporal backends; and pooling, viewing principled community detection as a theory-grounded coarse-graining operator for temporal graphs. Our results indicate that algorithmic methods remain the appropriate tool where attributes are absent or weak - scalability rather than accuracy being the binding obstacle - while neural models are most compelling when structural, temporal, and attribute signals align. By making temporal clustering scalable, GPU-accelerated primitives suggest a route toward theory-grounded pooling, while raising a central question: when does community-based coarse-graining preserve the dynamics needed for downstream learning tasks?
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。