揭示随机网络蒸馏与贝叶斯推断、深度集成的理论等价性
On the Equivalence of Random Network Distillation, Deep Ensembles, and Bayesian Inference
- 在无限宽网络极限下,通过神经正切核框架分析RND
- RND的自预测误差平方等价于深度集成的预测方差
- 构建贝叶斯RND模型可生成精确贝叶斯后验样本
不确定性量化是深度学习安全高效部署的核心,但许多计算实用方法缺乏严格的理论依据。随机网络蒸馏(RND)是一种轻量级技术,通过固定随机目标的预测误差来度量新奇性。尽管实证有效,其测量的不确定性类型及其与其他方法(如贝叶斯推断或深度集成)的关系仍不明确。本文在无限网络宽度极限下,基于神经正切核框架分析RND,发现两个核心结论:(1)RND的不确定性信号——其平方自预测误差——等价于深度集成的预测方差;(2)通过构造特定的RND目标函数,可使RND误差分布逼近宽神经网络贝叶斯推断的中心化后验预测分布。基于此等价性,进一步提出一种后验采样算法,利用该改进的贝叶斯RND模型生成独立同分布的精确贝叶斯后验预测样本。总体而言,研究为RND提供了统一的理论视角,将其置于深度集成与贝叶斯推断的严谨框架中,并开辟了高效且理论可靠的不确定性量化新路径。
原文摘要 · Abstract (English)
Uncertainty quantification is central to safe and efficient deployments of deep learning models, yet many computationally practical methods lack lacking rigorous theoretical motivation. Random network distillation (RND) is a lightweight technique that measures novelty via prediction errors against a fixed random target. While empirically effective, it has remained unclear what uncertainties RND measures and how its estimates relate to other approaches, e.g. Bayesian inference or deep ensembles. This paper establishes these missing theoretical connections by analyzing RND within the neural tangent kernel framework in the limit of infinite network width. Our analysis reveals two central findings in this limit: (1) The uncertainty signal from RND -- its squared self-predictive error -- is equivalent to the predictive variance of a deep ensemble. (2) By constructing a specific RND target function, we show that the RND error distribution can be made to mirror the centered posterior predictive distribution of Bayesian inference with wide neural networks. Based on this equivalence, we moreover devise a posterior sampling algorithm that generates i.i.d. samples from an exact Bayesian posterior predictive distribution using this modified \textit{Bayesian RND} model. Collectively, our findings provide a unified theoretical perspective that places RND within the principled frameworks of deep ensembles and Bayesian inference, and offer new avenues for efficient yet theoretically grounded uncertainty quantification methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。