arXiv:2511.17378cs.LG2025-11NeurIPS被引 3

揭示了数据一致性如何影响优化器选择简单解的机制

A Unified Stability Analysis of SAM vs SGD: Role of Data Coherence and Emergence of Simplicity Bias

  • 提出用数据一致性度量梯度曲率对齐程度
  • 发现一致性强的数据更易收敛到平坦解
  • 适用于研究优化器泛化能力的科研人员

理解深度学习中的优化动态在模型规模扩大背景下愈发重要。尽管随机梯度下降(SGD)及其变体能稳定找到泛化性能良好的解,但其背后的机制仍不明确。尤其在过参数化设置下,这些算法常偏好更平坦或更简单的极小值。已有研究将平坦性与泛化能力关联,而尖锐感知最小化(SAM)等方法则显式鼓励平坦性,但尚未建立统一理论来连接数据结构、优化动态与学习解的本质。本文构建了一个线性稳定性框架,分析了SGD、随机扰动和SAM在两层ReLU网络中的行为。分析核心是一个衡量梯度曲率在不同数据点间对齐程度的一致性指标,揭示了为何某些极小值在训练过程中更稳定且被优先选择。

原文摘要 · Abstract (English)

Understanding the dynamics of optimization in deep learning is increasingly important as models scale. While stochastic gradient descent (SGD) and its variants reliably find solutions that generalize well, the mechanisms driving this generalization remain unclear. Notably, these algorithms often prefer flatter or simpler minima, particularly in overparameterized settings. Prior work has linked flatness to generalization, and methods like Sharpness-Aware Minimization (SAM) explicitly encourage flatness, but a unified theory connecting data structure, optimization dynamics, and the nature of learned solutions is still lacking. In this work, we develop a linear stability framework that analyzes the behavior of SGD, random perturbations, and SAM, particularly in two layer ReLU networks. Central to our analysis is a coherence measure that quantifies how gradient curvature aligns across data points, revealing why certain minima are stable and favored during training.

优化器泛化能力稳定性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。