arXiv:2502.10162cs.LGcs.AI2025-02被引 2

从交互关系角度重新理解DNN泛化能力,发现可泛化交互呈衰减分布。

Revisiting Generalization Power of a DNN in Terms of Symbolic Interactions

  • 将DNN泛化能力归因于特征间交互的泛化性,提出交互分布分析新视角。
  • 可泛化交互服从衰减分布,非泛化交互呈纺锤形分布,实验证实匹配真实情况。
  • 适合关注模型内在机理、泛化理论或可解释性的研究者阅读。

本文从交互角度重新分析深度神经网络(DNN)的泛化能力。不同于以往在高维特征空间中的分析,我们发现DNN的泛化能力可归因于其交互项的泛化性。研究发现,可泛化的交互项遵循衰减型分布,而不可泛化的交互项则呈现纺锤形分布。此外,我们的理论能有效解耦这两类交互。实验表明,该理论能良好匹配DNN中真实存在的交互模式。

原文摘要 · Abstract (English)

This paper aims to analyze the generalization power of deep neural networks (DNNs) from the perspective of interactions. Unlike previous analysis of a DNN's generalization power in a highdimensional feature space, we find that the generalization power of a DNN can be explained as the generalization power of the interactions. We found that the generalizable interactions follow a decay-shaped distribution, while non-generalizable interactions follow a spindle-shaped distribution. Furthermore, our theory can effectively disentangle these two types of interactions from a DNN. We have verified that our theory can well match real interactions in a DNN in experiments.

泛化能力交互分析深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。