arXiv:2512.07988cs.LGcs.GR2025-12被引 2

用拓扑分析神经网络中间特征,揭示分类边界与鲁棒性规律

HOLE: Homological Observation of Latent Embeddings for Neural Network Interpretability

  • 通过持久同调提取中间激活的拓扑特征
  • 发现类别分离、特征解耦与模型鲁棒性的拓扑模式
  • 适合关注模型可解释性与结构分析的研究者

深度学习模型在多个领域取得显著成功,但其学习到的表征和决策过程仍高度不透明。本文提出HOLE(Homological Observation of Latent Embeddings),一种基于持久同调的分析方法,用于解析判别性神经网络的中间表征。HOLE从中间激活中提取拓扑特征,并通过簇流图、斑块图和热力图树状图等可视化工具呈现,帮助考察各层表征的结构与质量。我们在多种判别模型上评估HOLE,重点关注表征质量、跨层可解释性以及对输入扰动和模型压缩的鲁棒性。结果表明,拓扑分析能揭示与类别分离、特征解耦及模型鲁棒性相关的模式,为理解与改进深度学习系统提供了互补视角。

原文摘要 · Abstract (English)

Deep learning models have achieved remarkable success across various domains, yet their learned representations and decision-making processes remain largely opaque and hard to interpret. This work introduces HOLE (Homological Observation of Latent Embeddings), a method for analyzing and interpreting discriminative neural networks through persistent homology. HOLE extracts topological features from intermediate activations and presents them using a suite of visualization techniques, including cluster flow diagrams, blob graphs, and heatmap dendrograms. These tools facilitate the examination of representation structure and quality across layers. We evaluate HOLE using a range of discriminative models, focusing on representation quality, interpretability across layers, and robustness to input perturbations and model compression. The results indicate that topological analysis reveals patterns associated with class separation, feature disentanglement, and model robustness, providing a complementary perspective for understanding and improving deep learning systems.

可解释性拓扑分析神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。