arXiv:2511.04808cs.LG2025-11被引 3

数据越多,尖锐最小值越容易被找到并保持良好泛化

Sharp Minima Can Generalize: A Loss Landscape Perspective On Data

  • 通过分析损失曲面,发现数据量影响最小值体积分布
  • 小数据时尖锐最小值虽能泛化但体积太小难被找到
  • 大数据使原本尖锐的最小值相对变大,更易被优化器捕获

体积假说认为深度学习有效是因为其倾向于找到体积大的平坦最小值,而这些最小值具有良好的泛化能力。然而,这一观点无法解释大规模数据集在泛化中的作用。通过测量不同训练数据量下的最小值体积,发现尽管存在能够泛化的尖锐最小值,但由于其体积过小,难以被找到。随着数据量增加,损失曲面发生变化,使得原先体积小且能泛化的最小值变得(相对)更大,从而更容易被优化算法探索到。

原文摘要 · Abstract (English)

The volume hypothesis suggests deep learning is effective because it is likely to find flat minima due to their large volumes, and flat minima generalize well. This picture does not explain the role of large datasets in generalization. Measuring minima volumes under varying amounts of training data reveals sharp minima which generalize well exist, but are unlikely to be found due to their small volumes. Increasing data changes the loss landscape, such that previously small generalizing minima become (relatively) large.

损失曲面泛化能力数据规模最小值体积

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。