提出更优的子网络拉普拉斯近似方法,显著提升神经网络不确定性量化精度。
Optimality of Sub-network Laplace Approximations: New Results and Methods

- 基于梯度和精度矩阵交互设计新参数选择策略
- 证明现有方法系统低估预测方差,且保留越多参数偏差越小
- 适用于需可靠置信度估计的模型部署场景
尽管拉普拉斯近似为深度神经网络提供了简单的不确定性量化途径,但其对大型海森矩阵求逆的依赖催生了多种计算可行的低维或稀疏近似方法。其中代表性方法——子网络拉普拉斯近似,通过仅关注少量参数构建代理模型。现有方法通常依赖对角线、层内或其它架构启发式选择参数子集,忽略参数间相互作用,且缺乏形式化最优性保证。本文对子网络拉普拉斯范式进行了严格理论分析,证明所有此类方法均系统性低估全拉普拉斯后验的预测方差,且该偏差随保留子矩阵扩大而单调递减。基于此洞察,我们提出两种原理清晰、解析基础牢固的子网络海森近似方法: extit{Gradient-Laplace} 依据模型输出对参数的平均平方梯度在参考数据集上的大小选择参数; extit{Greedy-Laplace} 则通过考虑精度矩阵中的非对角交互,迭代优化选择。我们建立了刻画其最优性性质的理论保证,并证明 extit{Gradient-Laplace} 在理论上优于现有启发式方法。跨多种设置的大量数值实验表明,这些方法相对于现有基准表现优异。
原文摘要 · Abstract (English)
Although the Laplace approximation offers a simple route to uncertainty quantification in deep neural networks, its reliance on inverting large Hessian matrices has motivated a range of computationally feasible low-dimensional or sparse approximations. A prominent class of such methods - sub-network Laplace approximations, constructs surrogates by restricting attention to a small subset of parameters. Existing approaches in this family typically rely on diagonal, layer-wise, or other architectural heuristics for subset selection, which ignore cross-parameter interactions and lack formal optimality guarantees. In this paper, we provide a rigorous theoretical analysis of the sub-network Laplace paradigm. We prove that all sub-network Laplace methods systematically underestimate the predictive variance of the full Laplace posterior, and that this bias decreases monotonically as the retained sub-matrix expands. Leveraging this insight, we propose two principled, analytically grounded sub-network Hessian approximations: \textit{Gradient-Laplace} selects parameters with the largest average squared gradients of the model output with respect to the parameters over a reference dataset; while \textit{Greedy-Laplace} iteratively refines this selection by accounting for off-diagonal interactions in the precision matrix. We establish theoretical guarantees characterizing their optimality properties and show that Gradient-Laplace provably outperforms existing heuristic approaches. Extensive numerical studies across diverse settings indicate that these methods perform strongly relative to existing benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。