arXiv:2506.20779stat.MLcs.LG2025-06NeurIPS被引 6

高维下平坦解泛化能力骤降,因神经元‘碎裂’激活机制。

Stable Minima of ReLU Neural Networks Suffer from the Curse of Dimensionality: The Neural Shattering Phenomenon

  • 构建边界聚焦的ReLU神经元构造,揭示高维下平坦解的脆弱性
  • 证明平坦解的泛化误差随维度指数级恶化,速率低于低范数解
  • 首次系统解释平坦最小值在高维中失效的原因,适合关注泛化理论者

我们研究了两层过参数化ReLU网络中平坦性/低损失曲率的隐式偏差及其对泛化的影响,针对多变量输入这一更现实场景。现有工作或依赖插值,或仅限单变量输入。本文在两类自然设定下——(1) 平坦解的泛化差距,(2) 稳定最小值在非参数函数估计中的均方误差(MSE)——建立了上下界,表明尽管平坦性确实暗示泛化,但收敛速率随输入维度增长而指数级下降。这导致平坦解与已知不受维度诅咒影响的低范数解之间存在指数级差距。特别地,我们的极小极大下界构造基于一种新颖的打包论证,引入边界局部化的ReLU神经元,揭示了平坦解如何通过极少激活但权重极大的“神经碎裂”机制实现低曲率,却导致高维性能差。数值实验验证了这些理论发现。据我们所知,这是首个系统解释平坦最小值在高维中可能无法泛化的分析。

原文摘要 · Abstract (English)

We study the implicit bias of flatness / low (loss) curvature and its effects on generalization in two-layer overparameterized ReLU networks with multivariate inputs -- a problem well motivated by the minima stability and edge-of-stability phenomena in gradient-descent training. Existing work either requires interpolation or focuses only on univariate inputs. This paper presents new and somewhat surprising theoretical results for multivariate inputs. On two natural settings (1) generalization gap for flat solutions, and (2) mean-squared error (MSE) in nonparametric function estimation by stable minima, we prove upper and lower bounds, which establish that while flatness does imply generalization, the resulting rates of convergence necessarily deteriorate exponentially as the input dimension grows. This gives an exponential separation between the flat solutions compared to low-norm solutions (i.e., weight decay), which are known not to suffer from the curse of dimensionality. In particular, our minimax lower bound construction, based on a novel packing argument with boundary-localized ReLU neurons, reveals how flat solutions can exploit a kind of "neural shattering" where neurons rarely activate, but with high weight magnitudes. This leads to poor performance in high dimensions. We corroborate these theoretical findings with extensive numerical simulations. To the best of our knowledge, our analysis provides the first systematic explanation for why flat minima may fail to generalize in high dimensions.

泛化理论深度学习维度诅咒神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。