系统梳理深度学习不确定性量化方法,助模型在关键场景更可信
Uncertainty quantification for trustworthy deep learning: Methods and measures
- 区分生成预测分布的方法与总结不确定性的度量,结构清晰
- 对比五类方法,揭示高效集成近似与单次前向传播的性能优势
- 适合关注模型可信度、安全部署的研究者与工程师
深度神经网络在安全关键领域应用需可靠的预测置信度估计,但传统架构缺乏规范的不确定性量化。本文系统综述了基于集成与近似贝叶斯的不确定性量化(UQ)方法及其度量方式。聚焦于贝叶斯神经网络、蒙特卡洛丢弃、深度集成、高效集成近似及最后一层或单次前向方法五类,厘清其理论动机、实现与实证表现。整合证据网络、先验网络、分位数预测与事后校准等关联工作,涵盖分布外检测与选择性预测等决策任务。分析集成多样性理论与不确定性度量分解,对比熵分解与成对发散度,统一评估方法以支撑可比性。最后讨论大语言模型中的不确定性及开放方向:分类任务中高效的认知不确定性度量、最后一层多样性、分布偏移下的多样性与校准、混合架构设计。
原文摘要 · Abstract (English)
The deployment of deep neural networks in safety-critical domains demands reliable estimates of predictive confidence, yet conventional architectures lack principled uncertainty quantification. This survey provides a structured, critical review of methods for Uncertainty Quantification (UQ) in deep learning, scoped to ensemble-based and approximate Bayesian approaches and the measures used to summarize their outputs. Relative to existing UQ surveys, our contribution is depth on efficient ensemble approximations and single-pass methods, and a unified treatment that separates the method producing a predictive distribution from the measure that summarizes its uncertainty. We organize methods into five families: Bayesian neural networks, Monte Carlo Dropout, deep ensembles, efficient ensemble approximations, and last-layer or single-pass approaches. We situate adjacent work on evidential and prior networks, conformal prediction, and post-hoc calibration, together with the decision-time tasks of out-of-distribution detection and selective prediction. For each, we examine theoretical motivation, implementation, empirical performance, and limitations. We then review ensemble diversity theory and uncertainty measures and their decompositions, contrasting the entropy decomposition with pairwise divergence measures, and consolidate evaluation methodology so that our qualitative comparisons share a common basis. We close with a brief treatment of uncertainty in large language models and open research directions, including efficient epistemic measures for classification, last-layer diversity, diversity and calibration under shift, and hybrid architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。