arXiv:2605.05209cs.LGcs.AI2026-05被引 1

深度网络泛化能力与参数平坦性无关,自由度才是关键。

Are Flat Minima an Illusion?

  • 通过重缩放ReLU网络改变曲率,预测结果不变
  • 用特征自由度衡量泛化,相关性达0.47
  • 适合关注模型可适应性的研究者

平坦极小值常被用来解释深度网络的泛化能力,但平坦性是参数层面的描述,而泛化是函数层面的特性。同一函数可通过多种参数化实现。本文通过重缩放ReLU网络,使原始海森迹变化高达99倍,但所有预测保持不变,说明原始曲率无法揭示函数级解释。先前理论指出泛化源于函数所隐含约束的弱性,即模型在已学知识范围内保留的未来适应自由度。为量化此自由度,冻结最后一层隐藏表示,测试512个标签组合能否由替换的仿射分类器实现,计算联合完成得分。该指标对特征坐标的可逆线性变换和移位保持不变。在两组各100个预定义网络中,该得分对保留精度的秩相关系数分别为0.29和0.47。原始海森迹和相对平坦度均无经多重性校正后的关联。核心观点:自由度关联适应能力,平坦性仅为描述方式。

原文摘要 · Abstract (English)

Flat minima are an account of why deep networks generalise. However flatness is a matter of form (parameters), while generalisation is of function. The same function can be a result of many different parameterisations. I demonstrate this by rescaling ReLU networks, changing raw Hessian trace by up to $99$ times while every prediction remains fixed. Raw curvature cannot identify a function-level explanation. Previous theoretical work traced generalisation to the weakness of constraints implied by function, meaning the freedom a model retains within the bounds of what it has learned to be correct. A policy is weaker when more future commitments remain compatible with what it has learned, allowing more freedom to adapt. To measure this for neural networks, I freeze the last hidden representation and ask whether each of 512 sampled label bundles can be met by a replacement affine classifier. The resulting joint completion score is invariant under invertible linear mixing and translation of feature coordinates. Across two predeclared cohorts of 100 networks, it predicts held-out accuracy with rank correlations $0.29$ and $0.47$. Raw Hessian trace and relative flatness have no multiplicity-corrected association. To put it provocatively, freedom is correlated with adaptability, while flatness is a matter of description.

泛化能力模型自由度神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。