首次量化分析ReLU网络的几何折叠效应,揭示输入空间如何被非凸变形。
On Space Folds of ReLU Neural Networks
- 通过范围度量分析输入直线在激活空间的折叠行为
- 发现输入空间的直线在激活后普遍失去凸性,呈现非凸折叠
- 适用于对神经网络几何特性感兴趣的研究者
近期研究表明,ReLU神经网络的连续层可被理解为输入空间的几何折叠变换,展现出自相似模式。本文首次对这一空间折叠现象进行定量分析。我们聚焦于欧氏输入空间中的直线路径如何映射到汉明激活空间的对应路径。在此过程中,直线的凸性通常被破坏,导致非凸折叠行为。为此,我们引入一种基于范围度量的新指标,类似于随机游走研究中的方法,并证明了输入空间与激活空间中凸性概念的等价性。此外,我们在一个几何分析基准(CantorNet)和一个图像分类基准(MNIST)上进行了实证分析。本工作通过几何折叠现象深化了对ReLU神经网络激活空间的理解,为模型处理输入信息提供了重要洞见。
原文摘要 · Abstract (English)
Recent findings suggest that the consecutive layers of ReLU neural networks can be understood geometrically as space folding transformations of the input space, revealing patterns of self-similarity. In this paper, we present the first quantitative analysis of this space folding phenomenon in ReLU neural networks. Our approach focuses on examining how straight paths in the Euclidean input space are mapped to their counterparts in the Hamming activation space. In this process, the convexity of straight lines is generally lost, giving rise to non-convex folding behavior. To quantify this effect, we introduce a novel measure based on range metrics, similar to those used in the study of random walks, and provide the proof for the equivalence of convexity notions between the input and activation spaces. Furthermore, we provide empirical analysis on a geometrical analysis benchmark (CantorNet) as well as an image classification benchmark (MNIST). Our work advances the understanding of the activation space in ReLU neural networks by leveraging the phenomena of geometric folding, providing valuable insights on how these models process input information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。