揭示低温下吉布斯算法的泛化误差规律,支持平坦极小值有益于泛化。
Generalization of the Gibbs algorithm with high probability at low temperatures
- 基于先验体积推导低温度下的泛化误差上界
- 泛化误差随相似或更优经验误差假设的先验体积增大而减小
- 理论支持平坦极小值提升泛化能力,适用于多种随机算法
本文给出了吉布斯算法的泛化误差上界,该上界在高温区间恢复了已知的数据无关边界,并扩展至低温区间。在低温下,泛化性能关键取决于数据相关的损失曲面结构。研究表明,以高概率成立时,从吉布斯后验中抽取的单个假设的泛化误差会随着所有具有相似或更优经验误差的假设的总先验体积增加而减小。这为平坦极小值有益于泛化的观点提供了理论支持。文章还讨论了零温度极限,并将该边界推广至一类类似的随机算法。
原文摘要 · Abstract (English)
The paper gives a bound on the generalization error of the Gibbs algorithm, which recovers known data-independent bounds for the high temperature range and extends to the low-temperature range, where generalization depends critically on the data-dependent loss-landscape. It is shown, that with high probability the generalization error of a single hypothesis drawn from the Gibbs posterior decreases with the total prior volume of all hypotheses with similar or smaller empirical error. This gives theoretical support to the belief in the benefit of flat minima. The zero temperature limit is discussed and the bound is extended to a class of similar stochastic algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。