用随机网络对大模型输出做后处理,精准估算预测不确定性
Uncertainty Quantification for Large-Scale Deep Networks via Post-StoNet Modeling
- 将大模型最后一层输出输入随机神经网络,加稀疏正则训练
- 预测区间更短且更诚实,校准效果优于同类方法
- 首次将线性模型稀疏学习理论拓展到深度神经网络
深度学习已重塑现代数据科学,但大规模深度神经网络(DNN)预测的不确定性量化仍是个未解难题。为此,我们提出一种新型后处理方法:将预训练大模型最后一层隐藏层输出输入随机神经网络(StoNet),在验证集上施加稀疏惩罚进行训练,并构建未来观测的预测区间。我们建立了该方法的有效性理论保证;其中,稀疏StoNet的参数估计一致性是关键。大量实验表明,该方法可生成更短且更诚实的置信区间,优于传统合规方法;同时在校准性能上也超越其他后处理校准技术。此外,StoNet框架为将线性模型中的稀疏学习理论与方法迁移至DNN提供了平台。
原文摘要 · Abstract (English)
Deep learning has revolutionized modern data science. However, how to accurately quantify the uncertainty of predictions from large-scale deep neural networks (DNNs) remains an unresolved issue. To address this issue, we introduce a novel post-processing approach. This approach feeds the output from the last hidden layer of a pre-trained large-scale DNN model into a stochastic neural network (StoNet), then trains the StoNet with a sparse penalty on a validation dataset and constructs prediction intervals for future observations. We establish a theoretical guarantee for the validity of this approach; in particular, the parameter estimation consistency for the sparse StoNet is essential for the success of this approach. Comprehensive experiments demonstrate that the proposed approach can construct honest confidence intervals with shorter interval lengths compared to conformal methods and achieves better calibration compared to other post-hoc calibration techniques. Additionally, we show that the StoNet formulation provides us with a platform to adapt sparse learning theory and methods from linear models to DNNs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。