提出新基准评估神经网络局部采样,发现RMSProp预条件SGLD最有效。
From Global to Local: A Scalable Benchmark for Local Posterior Sampling
- 从全局采样转向局部采样,构建可扩展的评估基准
- 在百万级参数模型中成功提取非平凡局部信息
- 适用于大规模模型后验分布分析,尤其关注局部几何
神经网络损失函数的退化性是固有特征,但现有随机梯度马尔可夫链蒙特卡洛(SGMCMC)算法如何与之交互尚不明确。常见SGMCMC算法的全局收敛性理论依赖于与退化损失景观不兼容的假设。本文主张将研究重点从全局转向局部后验采样,并首次提出一个可扩展的基准来评估SGMCMC算法的局部采样性能。我们评估了多种算法,发现RMSProp预条件SGLD在忠实表征后验分布局部几何方面表现最优。尽管缺乏全局收敛的理论保证,实验结果表明,该方法可在高达O(100M)参数的模型中提取有意义的局部信息。
原文摘要 · Abstract (English)
Degeneracy is an inherent feature of the loss landscape of neural networks, but it is not well understood how stochastic gradient MCMC (SGMCMC) algorithms interact with this degeneracy. In particular, existing global convergence guarantees for common SGMCMC algorithms rely on assumptions which are likely incompatible with degenerate loss landscapes. In this paper, we argue that this gap requires a shift in focus from global to local posterior sampling, and, as a first step, we introduce a novel scalable benchmark for evaluating the local sampling performance of SGMCMC algorithms. We evaluate a number of common algorithms, and find that RMSProp-preconditioned SGLD is most effective at faithfully representing the local geometry of the posterior distribution among the samplers we evaluate. Although we lack theoretical guarantees about global sampler convergence, our empirical results show that we are able to extract non-trivial local information in models with up to O(100M) parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。