arXiv:2601.21835cs.LG2026-01中稿 · European Symposium…被引 2

用神经核替代计算量大的雅可比矩阵,实现大模型的高效不确定性估计。

Scalable Linearized Laplace Approximation via Surrogate Neural Kernel

  • 用代理网络学习紧凑特征,通过内积模拟神经正切核
  • 在大规模预训练模型上实现与现有方法相当或更好的不确定性校准
  • 优化核函数能显著提升分布外样本检测能力,适合需要可信预测的场景

我们提出一种可扩展的方法来近似线性化拉普拉斯近似(LLA)的核函数。该方法使用一个代理深度神经网络(DNN),学习一种紧凑的特征表示,其内积能复现神经正切核(NTK),从而避免计算大型雅可比矩阵。训练仅依赖高效的雅可比-向量乘积,可在大规模预训练DNN上计算预测不确定性。实验表明,该方法在不确定性估计和校准方面达到或超过现有LLA近似效果。更重要的是,对学习到的核函数进行偏差调整能显著提升分布外检测性能。这表明,在给定预训练DNN的前提下,该方法可通过设计更优核函数来改进预测不确定性计算。

原文摘要 · Abstract (English)

We introduce a scalable method to approximate the kernel of the Linearized Laplace Approximation (LLA). For this, we use a surrogate deep neural network (DNN) that learns a compact feature representation whose inner product replicates the Neural Tangent Kernel (NTK). This avoids the need to compute large Jacobians. Training relies solely on efficient Jacobian-vector products, allowing to compute predictive uncertainty on large-scale pre-trained DNNs. Experimental results show similar or improved uncertainty estimation and calibration compared to existing LLA approximations. Notwithstanding, biasing the learned kernel significantly enhances out-of-distribution detection. This remarks the benefits of the proposed method for finding better kernels than the NTK in the context of LLA to compute prediction uncertainty given a pre-trained DNN.

不确定性估计神经核大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。