研究网络谱嵌入中度归一化的最优选择,揭示其在不同网络条件下的效果差异。
Spectral Embeddings of Degree-$α$ Laplacians in Random Dot Product Graphs

- 提出一族度归一化谱嵌入,涵盖邻接矩阵和对称拉普拉斯等常见形式。
- 在随机点积图模型下证明嵌入的中心极限定理,揭示归一化对节点分布的影响。
- 发现无统一最优归一化方式,低密度或不平衡网络更需强归一化。
针对网络数据的谱聚类方法通常基于邻接矩阵或对称拉普拉斯矩阵等矩阵表示。本文研究了一族包含这些常见形式作为特例的度归一化谱嵌入。在随机点积图模型下,建立了该族嵌入的逐行中心极限定理。结果明确描述了度归一化如何影响总体几何结构和嵌入节点的局部不确定性。利用极限分布,通过投影高斯贝叶斯误差诊断,在两社区随机块模型中比较不同归一化方式。结果显示,不存在单一归一化始终更优;优选归一化取决于网络密度、社区不平衡度及块概率结构。通常在低密度或高度不平衡设置下,更强的归一化更受青睐。这些结果为替代归一化在何时何地能提升谱聚类提供了统一的概率解释。
原文摘要 · Abstract (English)
Spectral clustering methods for network data are commonly based on a few matrix representations, such as the adjacency matrix and the symmetric Laplacian. We study a continuum of degree-normalized spectral embeddings that includes these commonly used choices as special cases. Under a random dot product graph model, we establish a row-wise central limit theorem for this family of embeddings. The result provides an explicit description of how degree normalization affects both population geometry and the local uncertainty of embedded nodes. We use the limiting distributions to compare different normalizations in two-community stochastic block models through a projected-Gaussian Bayes-error diagnostic. These comparisons show that no single normalization is uniformly preferred. Instead, the favored normalization depends on network density, community imbalance, and block-probability structure. Typically, stronger normalization is favored in lower-density or more imbalanced settings. These results provide a unified distributional understanding of when and why alternative normalizations may improve spectral clustering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。