arXiv:2410.16901cs.LGstat.ML2024-10被引 4

让贝叶斯深度学习不欠拟合,保持预测精度的同时量化不确定性。

Bayes without Underfitting: Fully Correlated Deep Learning Posteriors via Alternating Projections

  • 在参数的零空间中构建贝叶斯近似,确保预测不劣于点估计。
  • 算法可线性扩展至2800万参数的视觉变换模型,保持高效计算。
  • 适合需要高置信度预测与准确不确定性的大型生成模型应用。

贝叶斯深度学习常因欠拟合导致预测精度低于点估计,使不确定性量化以牺牲准确性为代价。对于线性化模型,广义高斯-牛顿矩阵的零空间对应于不改变点估计训练预测的参数方向。本文提出在该零空间内构建贝叶斯近似,从而保证贝叶斯预测不会欠拟合。我们设计了一种无矩阵投影算法,其计算复杂度随参数量线性增长,随输出维度二次增长。进一步提出一种仅随参数量线性扩展的近似方法,使该方法适用于生成模型。大量实验证明,该方法可扩展至大型模型,包括含2800万参数的视觉变换模型。

原文摘要 · Abstract (English)

Bayesian deep learning all too often underfits so that the Bayesian prediction is less accurate than a simple point estimate. Uncertainty quantification then comes at the cost of accuracy. For linearized models, the null space of the generalized Gauss-Newton matrix corresponds to parameters that preserve the training predictions of the point estimate. We propose to build Bayesian approximations in this null space, thereby guaranteeing that the Bayesian predictive does not underfit. We suggest a matrix-free algorithm for projecting onto this null space, which scales linearly with the number of parameters and quadratically with the number of output dimensions. We further propose an approximation that only scales linearly with parameters to make the method applicable to generative models. An extensive empirical evaluation shows that the approach scales to large models, including vision transformers with 28 million parameters.

贝叶斯深度学习不确定性量化零空间投影大规模模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。