模型最优解的平坦与否,本质取决于所学函数复杂度。
A Function-Centric Perspective on Flat and Sharp Minima
- 从函数视角看极小值,平坦性是相对的,不绝对代表泛化好坏。
- 正则化后更尖锐的极小值反而提升泛化、校准和鲁棒性。
- 适合关注模型几何与泛化关系的研究者阅读。
深度神经网络中,平坦极小值常被认为有助于提升泛化性能。然而近年研究发现这一关联并不简单,存在理论反例与实证例外。本文重新审视尖锐性在模型表现中的作用,提出尖锐性应被理解为函数相关的属性,而非泛化差的指标。通过涵盖单目标优化、合成非线性二分类任务以及现代图像分类任务的广泛实验,我们发现:在单目标优化中,不同最优解具有显著不同的局部几何;在合成任务中,决策边界越紧致,尖锐性越高,但模型仍能完美泛化,说明尖锐性不等于记忆;在大规模实验中,使用权重衰减、数据增强或 SAM 等正则化方法时,更尖锐的极小值往往伴随更好的泛化、校准、鲁棒性和功能一致性。结果表明,函数复杂度决定了解空间几何,尖锐极小值可能反映更合适的归纳偏置,呼吁从函数中心视角重审极小值几何。
原文摘要 · Abstract (English)
Flat minima are strongly associated with improved generalisation in deep neural networks. However, this connection has proven nuanced in recent studies, with both theoretical counterexamples and empirical exceptions emerging in the literature. In this paper, we revisit the role of sharpness in model performance and argue that sharpness is better understood as a function-dependent property rather than an indicator of poor generalisation. We conduct extensive empirical studies ranging from single-objective optimisation, synthetic non-linear binary classification tasks, to modern image classification tasks. In single-objective optimisation, we show that flatness and sharpness are relative to the function being learned: equally optimal solutions can exhibit markedly different local geometry. In synthetic non-linear binary classification tasks, we show that increasing decision-boundary tightness can increase sharpness even when models generalise perfectly, indicating that sharpness is not reducible to memorisation alone. Finally, in large-scale experiments, we find that sharper minima often emerge when models are regularised (e.g., via weight decay, data augmentation, or SAM), and coincide with better generalisation, calibration, robustness, and functional consistency. Our findings suggest that function complexity, rather than flatness, shapes the geometry of solutions, and that sharper minima can reflect more appropriate inductive biases, calling for a function-centric reappraisal of minima geometry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。