提出对称拓扑度量,让神经网络表征分析更准确、可比。
Symmetric Divergence and Normalized Similarity: A Unified Topological Framework for Representation Analysis
- 引入对称拓扑差异度量SRTD,解决传统方法的不对称问题。
- 提出归一化拓扑相似性NTS,得分在-1到1间且与样本量无关。
- 适合研究模型结构变化或跨场景对比的学者使用。
拓扑数据分析(TDA)为比较神经网络表征提供了原理严谨的内在视角。然而,现有配对拓扑差异度量(如RTD)受限于启发式非对称性,且得分无界,依赖样本量,难以实现跨场景可靠基准测试。为此,我们构建了一个统一的拓扑工具包,满足精细结构诊断与鲁棒标准化评估双重需求。首先,通过引入对称表示拓扑差异(SRTD)及其高效变体SRTD-lite,完善了RTD框架。SRTD克服了以往方法的理论非对称性,将诊断信息整合为单一、全面的交叉条码签名,可精确定位结构差异,并作为有效优化目标,避免双方向计算开销。其次,为实现异构设置间的可靠基准测试,提出归一化拓扑相似性(NTS),通过测量层次合并顺序的秩相关性,获得范围在-1到1之间的尺度不变度量,有效克服未归一化差异度量的尺度与样本依赖问题。在合成与真实深度学习场景下的实验表明,该工具包能捕捉几何度量遗漏的功能性变化,并在距离饱和下仍稳健映射大语言模型谱系,提供一种严谨、拓扑感知的视角,补充诸如CKA等已有度量。
原文摘要 · Abstract (English)
Topological Data Analysis (TDA) offers a principled, intrinsic lens for comparing neural representations. However, existing paired topological divergences (e.g., RTD) are limited by heuristic asymmetry and, more critically, unbounded scores that depend on sample size, hindering reliable cross-scenario benchmarking. To address these challenges, we develop a unified topological toolkit serving two complementary needs: fine-grained structural diagnosis and robust, standardized evaluation. First, we complete the RTD framework by introducing Symmetric Representation Topology Divergence (SRTD) and its efficient variant SRTD-lite. Beyond resolving the theoretical asymmetry of prior variants, SRTD consolidates diagnostic information into a single, comprehensive cross-barcode signature. This allows for precise localization of structural discrepancies and serves as an effective optimization objective without the overhead of dual directional computations. Second, to enable reliable benchmarking across heterogeneous settings, we propose Normalized Topological Similarity (NTS). By measuring the rank correlation of hierarchical merge orders, NTS yields a scale-invariant metric bounded between -1 and 1, effectively overcoming the scale and sample-dependence of unnormalized divergences. Experiments across synthetic and real-world deep learning settings demonstrate that our toolkit captures functional shifts in CNNs missed by geometric measures and robustly maps LLM genealogy even under distance saturation, offering a rigorous, topology-aware perspective that complements measures like CKA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。