从交互关系角度重新理解DNN泛化能力,发现可泛化交互呈衰减分布。
Revisiting Generalization Power of a DNN in Terms of Symbolic Interactions
- 将DNN泛化能力归因于特征间交互的泛化性,提出交互分布分析新视角。
- 可泛化交互服从衰减分布,非泛化交互呈纺锤形分布,实验证实匹配真实情况。
- 适合关注模型内在机理、泛化理论或可解释性的研究者阅读。
本文从交互角度重新分析深度神经网络(DNN)的泛化能力。不同于以往在高维特征空间中的分析,我们发现DNN的泛化能力可归因于其交互项的泛化性。研究发现,可泛化的交互项遵循衰减型分布,而不可泛化的交互项则呈现纺锤形分布。此外,我们的理论能有效解耦这两类交互。实验表明,该理论能良好匹配DNN中真实存在的交互模式。
原文摘要 · Abstract (English)
This paper aims to analyze the generalization power of deep neural networks (DNNs) from the perspective of interactions. Unlike previous analysis of a DNN's generalization power in a highdimensional feature space, we find that the generalization power of a DNN can be explained as the generalization power of the interactions. We found that the generalizable interactions follow a decay-shaped distribution, while non-generalizable interactions follow a spindle-shaped distribution. Furthermore, our theory can effectively disentangle these two types of interactions from a DNN. We have verified that our theory can well match real interactions in a DNN in experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。