通过稀疏先验与变分推断,提升神经网络的鲁棒性与不确定性估计能力。
Variational Bayesian Bow tie Neural Networks with Shrinkage
- 引入稀疏先验和多层权重正则化,增强对网络结构变化的鲁棒性。
- 利用Polya-Gamma数据扩展实现快速近似推断,避免分布假设限制。
- 适用于需要高可靠性与不确定性建模的场景,如医疗或自动驾驶。
尽管深度模型在机器学习中占据主导地位,仍存在预测过于自信、易受对抗攻击及低估预测变异性等问题。贝叶斯范式为解决这些问题提供了自然框架,已成为深度模型不确定性估计的标准方法,并可提升精度与超参数调优能力。然而,精确贝叶斯推断困难,通常依赖于强独立性与分布假设的变分算法。此外,现有方法对网络架构选择敏感。本文针对标准前馈激活神经网络的随机松弛形式,采用权重上的稀疏促进先验,提升对架构设计的鲁棒性。借助Polya-Gamma数据增强技巧,将模型转化为条件线性和高斯形式,从而推导出一种快速近似变分推断算法,避免了分布假设与层间独立性限制。同时探讨了进一步提升可扩展性与处理多模态性的策略。
原文摘要 · Abstract (English)
Despite the dominant role of deep models in machine learning, limitations persist, including overconfident predictions, susceptibility to adversarial attacks, and underestimation of variability in predictions. The Bayesian paradigm provides a natural framework to overcome such issues and has become the gold standard for uncertainty estimation with deep models, also providing improved accuracy and a framework for tuning critical hyperparameters. However, exact Bayesian inference is challenging, typically involving variational algorithms that impose strong independence and distributional assumptions. Moreover, existing methods are sensitive to the architectural choice of the network. We address these issues by focusing on a stochastic relaxation of the standard feed-forward rectified neural network and using sparsity-promoting priors on the weights of the neural network for increased robustness to architectural design. Thanks to Polya-Gamma data augmentation tricks, which render a conditionally linear and Gaussian model, we derive a fast, approximate variational inference algorithm that avoids distributional assumptions and independence across layers. Suitable strategies to further improve scalability and account for multimodality are considered.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。