发现深度学习不确定性也遵循可预测的缩放规律。
Scaling Laws for Uncertainty in Deep Learning
- 通过实验验证不确定性随数据和模型规模变化的规律性
- 即使数据量大,认知不确定性仍不可忽略
- 为贝叶斯方法在大数据场景下的必要性提供证据
深度学习近年揭示了性能随数据集和模型规模变化的缩放规律。受此启发,我们探究类似规律是否适用于预测不确定性。在可识别参数模型中,通过贝叶斯视角可推导出不确定性缩放规律,例如认知不确定性以 $O(1/N)$ 速率随数据量 $N$ 收缩。但在过参数化模型中,此类保证不成立,行为尚不明确。本文通过视觉与语言任务的实验,首次实证发现多种不确定性度量(包括分布内与分布外)均随数据和模型规模呈现可预测的缩放规律。这些规律不仅具有理论美感,且可用于外推更大规模数据或模型下的不确定性。研究结果有力回应了对贝叶斯方法的常见质疑:‘数据这么多,为何还需要贝叶斯?’ 结果表明,‘数据多’并不足以使认知不确定性趋于零。
原文摘要 · Abstract (English)
Deep learning has recently revealed the existence of scaling laws, demonstrating that model performance follows predictable trends based on dataset and model sizes. Inspired by these findings and fascinating phenomena emerging in the over-parameterized regime, we examine a parallel direction: do similar scaling laws govern predictive uncertainties in deep learning? In identifiable parametric models, such scaling laws can be derived in a straightforward manner by treating model parameters in a Bayesian way. In this case, for example, we obtain $O(1/N)$ contraction rates for epistemic uncertainty with respect to the number of data $N$. However, in over-parameterized models, these guarantees do not hold, leading to largely unexplored behaviors. In this work, we empirically show the existence of scaling laws associated with various measures of predictive uncertainty with respect to dataset and model sizes. Through experiments on vision and language tasks, we observe such scaling laws for in- and out-of-distribution predictive uncertainty estimated through popular approximate Bayesian inference and ensemble methods. Besides the elegance of scaling laws and the practical utility of extrapolating uncertainties to larger data or models, this work provides strong evidence to dispel recurring skepticism against Bayesian approaches: "In many applications of deep learning we have so much data available: what do we need Bayes for?". Our findings show that "so much data" is typically not enough to make epistemic uncertainty negligible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。