arXiv:2606.29043cs.LG2026-06

探究尖锐度与复杂度联合解释模型泛化能力的边界

How Far Can Sharpness and Complexity Jointly Explain Generalization?

论文配图:How Far Can Sharpness and Complexity Jointly Explain Generalization?
图 1 · 摘自论文原文
  • 提出面向函数空间的新定义,弱化对参数表示依赖
  • 实证显示联合解释力优于单一指标,覆盖更广场景
  • 适用于理解模型泛化机制的研究者,但非完备理论

尖锐度与复杂度是深度神经网络泛化分析中的两大核心因素。现有对泛化度量的定量评估多聚焦于单一标量指标,尚未深入探索尖锐度与复杂度的联合解释能力。本文通过线性回归与基于帕累托的分析方法,定量评估二者联合解释泛化的能力。在现有参数级定义基础上,进一步提出更贴近函数空间、较少依赖原始参数表示的尖锐度与复杂度实现方式。结果表明,函数导向的定义显著拓展了两因子视角的解释范围,优于现有参数级度量。整体上,研究支持尖锐度-复杂度视角作为理解多种设置下泛化行为的有用工具;然而仍存在解释失败,表明该双因子视角是否能成为泛化理论的完整框架仍待验证。

原文摘要 · Abstract (English)

Sharpness and complexity are two central factors in the generalization analysis of deep neural networks. Existing quantitative evaluations of generalization measures have largely focused on individual scalar measures, leaving the joint explanatory power of sharpness and complexity largely unexplored. This work studies how far sharpness and complexity can jointly explain generalization. We use linear regression and introduce a Pareto-based analysis to quantitatively evaluate the joint explanatory power of these two factors. Beyond the existing parameter-level definitions, we further propose realizations of sharpness and complexity that are closer to function space and less dependent on raw parameter representations. We find that function-oriented definitions of these two quantities expand the explanatory scope of the two-factor view beyond what is achieved by existing parameter-level metrics. Overall, our results support the sharpness-complexity perspective as an informative lens for understanding generalization across diverse settings. At the same time, the remaining failures indicate that whether this two-factor view can serve as a complete theory of generalization remains open.

泛化分析尖锐度复杂度函数空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。