用瑞尼熵给出模型泛化误差上限,解释大模型为何不崩溃。
Overfitting has a limitation: a model-independent generalization gap bound based on Rényi entropy
- 基于数据直方图的算法泛化误差有新上界,仅依赖数据分布熵。
- 数据熵越低,大模型越不易过拟合,噪声会因提升熵导致性能下降。
- 为学习所需数据量提供理论下限,适合研究泛化与数据效率者。
机器学习模型持续扩大是否仍能成功?关键在于理解泛化误差(过拟合影响)。传统分析将误差界与模型复杂度关联,难以解释超大规模模型的成功。本文提出新视角:对仅由数据直方图决定输出的算法(如经验风险最小化或梯度方法),建立模型无关的泛化误差上界,该上界仅依赖于数据生成分布的瑞尼熵。结果表明,只要数据量足够,即使模型无限大,也能保持小的泛化误差。该框架可直接解释向数据注入随机噪声后性能显著下降的现象——因噪声增加了数据分布的瑞尼熵。此外,本文将无免费午餐定理推广为依赖数据分布的形式,证明学习所需的最小数据量恰等于瑞尼熵,凸显所提上界的紧致性。
原文摘要 · Abstract (English)
Will further scaling up of machine learning models continue to bring success? A significant challenge in answering this question lies in understanding generalization gap, which is the impact of overfitting. Understanding generalization gap behavior of increasingly large-scale machine learning models remains a significant area of investigation, as conventional analyses often link error bounds to model complexity, failing to fully explain the success of extremely large architectures. This research introduces a novel perspective by establishing a model-independent upper bound for generalization gap applicable to algorithms whose outputs are determined solely by the data's histogram, such as empirical risk minimization or gradient-based methods. Crucially, this bound is shown to depend only on the Rényi entropy of the data-generating distribution, suggesting that a small generalization gap can be maintained even with arbitrarily large models, provided the data quantity is sufficient relative to this entropy. This framework offers a direct explanation for the phenomenon where generalization performance degrades significantly upon injecting random noise into data, where the performance degrade is attributed to the consequent increase in the data distribution's Rényi entropy. Furthermore, we adapt the no-free-lunch theorem to be data-distribution-dependent, demonstrating that an amount of data corresponding to the Rényi entropy is indeed essential for successful learning, thereby highlighting the tightness of our proposed generalization bound.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。