用信息几何视角重新理解变分贝叶斯,揭示自然梯度的核心作用。
Information Geometry of Variational Bayes
- 将变分贝叶斯视为自然梯度计算,统一优化框架。
- 证明贝叶斯更新可简化为自然梯度相加,提升推断效率。
- 适用于大模型的高效变分推断,适合机器学习研究者。
我们揭示了信息几何与变分贝叶斯(VB)之间的一个基本联系,并探讨其对机器学习的影响。在特定条件下,变分贝叶斯解始终需要估计或计算自然梯度。通过使用Khan和Rue(2023)提出的自然梯度下降算法——贝叶斯学习规则(BLR),我们展示了这一事实的多个后果:(i)将贝叶斯法则简化为自然梯度的叠加;(ii)推广了梯度方法中常用的二次近似;(iii)实现了大规模语言模型的变分贝叶斯算法部署。尽管该联系及其后果并非全新,但我们进一步强调信息几何与贝叶斯方法的共同起源,旨在推动两领域交叉研究的发展。
原文摘要 · Abstract (English)
We highlight a fundamental connection between information geometry and variational Bayes (VB) and discuss its consequences for machine learning. Under certain conditions, a VB solution always requires estimation or computation of natural gradients. We show several consequences of this fact by using the natural-gradient descent algorithm of Khan and Rue (2023) called the Bayesian Learning Rule (BLR). These include (i) a simplification of Bayes' rule as addition of natural gradients, (ii) a generalization of quadratic surrogates used in gradient-based methods, and (iii) a large-scale implementation of VB algorithms for large language models. Neither the connection nor its consequences are new but we further emphasize the common origins of the two fields of information geometry and Bayes with a hope to facilitate more work at the intersection of the two fields.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。