arXiv:2602.01136cs.LGmath.DS2026-02

用谱分析统一量化深度模型的稳定性和可解释性。

A Unified Matrix-Spectral Framework for Stability and Interpretability in Deep Learning

  • 将网络视为线性算子的乘积,通过谱特征分析敏感性。
  • 引入全局矩阵稳定性指数,显著提升特征归因稳定性。
  • 适合关注模型鲁棒性与可解释性设计的研究者。

我们提出一种统一的矩阵-谱框架,用于分析深度神经网络的稳定性和可解释性。将网络表示为依赖数据的线性算子乘积,揭示了控制输入扰动、标签噪声和训练动态敏感性的谱量。我们引入全局矩阵稳定性指数(Global Matrix Stability Index),整合雅可比、参数梯度、神经正切核算子和损失赫斯蒂安的谱信息,形成一个控制前向敏感性、归因鲁棒性和优化条件的单一稳定性尺度。进一步表明,谱熵通过捕捉典型而非纯最坏情况的敏感性,改进了经典算子范数界。这些量提供可计算的诊断工具和面向稳定性的正则化原则。在MNIST、CIFAR-10和CIFAR-100上的合成实验与受控研究证实,适度的谱正则化能显著提升归因稳定性,即使全局谱摘要变化不大。结果建立了谱集中度与解析稳定性之间的精确联系,为感知鲁棒性的模型设计与训练提供实用指导。

原文摘要 · Abstract (English)

We develop a unified matrix-spectral framework for analyzing stability and interpretability in deep neural networks. Representing networks as data-dependent products of linear operators reveals spectral quantities governing sensitivity to input perturbations, label noise, and training dynamics. We introduce a Global Matrix Stability Index that aggregates spectral information from Jacobians, parameter gradients, Neural Tangent Kernel operators, and loss Hessians into a single stability scale controlling forward sensitivity, attribution robustness, and optimization conditioning. We further show that spectral entropy refines classical operator-norm bounds by capturing typical, rather than purely worst-case, sensitivity. These quantities yield computable diagnostics and stability-oriented regularization principles. Synthetic experiments and controlled studies on MNIST, CIFAR-10, and CIFAR-100 confirm that modest spectral regularization substantially improves attribution stability even when global spectral summaries change little. The results establish a precise connection between spectral concentration and analytic stability, providing practical guidance for robustness-aware model design and training.

深度学习稳定性可解释性谱分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。