让神经网络像高斯过程一样可靠,关键在可学习的激活函数。
Revisiting the Equivalence of Bayesian Neural Networks and Gaussian Processes: On the Importance of Learning Activations
- 用可学习激活函数实现贝叶斯网络对高斯过程的精确模拟
- 在多个数据集上优于已有方法,且理论基础更扎实
- 适合关注不确定性建模与模型可解释性的研究者
高斯过程(GPs)提供便捷的函数空间先验,是建模不确定性的自然选择。而贝叶斯神经网络(BNNs)虽具更好扩展性与可扩展性,却缺乏GP的优越性质。这促使人们发展能复现GP行为的BNN。然而现有方法或局限于特定核函数,或依赖启发式设计。本文证明,可训练激活函数对将GP先验映射至宽BNN至关重要。我们利用闭式2-Wasserstein距离,实现重参数化先验与激活函数的梯度优化。此外,引入可训练周期性激活函数,确保全局平稳性;设计基于GP超参数的条件函数先验,支持高效模型选择。实验表明,该方法在多个数据集上持续优于现有方法,或媲美启发式方法,同时具备更强理论支撑。
原文摘要 · Abstract (English)
Gaussian Processes (GPs) provide a convenient framework for specifying function-space priors, making them a natural choice for modeling uncertainty. In contrast, Bayesian Neural Networks (BNNs) offer greater scalability and extendability but lack the advantageous properties of GPs. This motivates the development of BNNs capable of replicating GP-like behavior. However, existing solutions are either limited to specific GP kernels or rely on heuristics. We demonstrate that trainable activations are crucial for effective mapping of GP priors to wide BNNs. Specifically, we leverage the closed-form 2-Wasserstein distance for efficient gradient-based optimization of reparameterized priors and activations. Beyond learned activations, we also introduce trainable periodic activations that ensure global stationarity by design, and functional priors conditioned on GP hyperparameters to allow efficient model selection. Empirically, our method consistently outperforms existing approaches or matches performance of the heuristic methods, while offering stronger theoretical foundations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。