arXiv:2510.21582cs.LG2025-10

揭示深度神经网络隐藏层中语义结构的生成机制。

An unsupervised tour through the hidden pathways of deep neural networks

  • 用无监督方法分析隐藏层表示的内在维度与概率分布变化。
  • 发现早期层消除无关结构,后期层形成类概念层次的密度峰值。
  • 解释了宽网络为何能泛化:冗余参数防止过拟合,需零训练误差与正则化。

本论文旨在提升对深度神经网络如何生成有意义表征并实现泛化的内部机制的理解。重点使用我们提出的无监督学习工具,刻画隐藏表示中的语义内容,这些工具利用数据的低维结构。第二章提出Gride方法,可在不删减数据集的情况下,以尺度为变量显式估计数据的内在维度;该方法基于严格的分布理论,可量化估计不确定性,且仅依赖最近邻点间距离,计算高效。第三章研究了先进深度神经网络中隐藏层概率密度的演化:初始层生成单峰分布,消除分类无关结构;后续层以层次化方式出现密度峰,反映概念的语义层级;输出层的峰拓扑结构可重构类别间的语义关系。第四章探讨泛化问题:在插值训练数据的网络中增加参数通常提升泛化性能,与经典偏差-方差权衡矛盾。我们证明,宽网络学习冗余表示而非拟合虚假相关性;冗余神经元仅在模型正则化且训练误差为零时出现。

原文摘要 · Abstract (English)

The goal of this thesis is to improve our understanding of the internal mechanisms by which deep artificial neural networks create meaningful representations and are able to generalize. We focus on the challenge of characterizing the semantic content of the hidden representations with unsupervised learning tools, partially developed by us and described in this thesis, which allow harnessing the low-dimensional structure of the data. Chapter 2. introduces Gride, a method that allows estimating the intrinsic dimension of the data as an explicit function of the scale without performing any decimation of the data set. Our approach is based on rigorous distributional results that enable the quantification of uncertainty of the estimates. Moreover, our method is simple and computationally efficient since it relies only on the distances among nearest data points. In Chapter 3, we study the evolution of the probability density across the hidden layers in some state-of-the-art deep neural networks. We find that the initial layers generate a unimodal probability density getting rid of any structure irrelevant to classification. In subsequent layers, density peaks arise in a hierarchical fashion that mirrors the semantic hierarchy of the concepts. This process leaves a footprint in the probability density of the output layer, where the topography of the peaks allows reconstructing the semantic relationships of the categories. In Chapter 4, we study the problem of generalization in deep neural networks: adding parameters to a network that interpolates its training data will typically improve its generalization performance, at odds with the classical bias-variance trade-off. We show that wide neural networks learn redundant representations instead of overfitting to spurious correlation and that redundant neurons appear only if the network is regularized and the training error is zero.

神经网络无监督学习泛化表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。