arXiv:2410.05757stat.MLcs.LG2024-10中稿 · UAI 2025被引 2

提出数据驱动方法自动优化贝叶斯深度学习的温度参数,提升预测性能。

Temperature Optimization for Bayesian Deep Learning

  • 将温度视为可学习参数,直接从数据中估计最优值。
  • 在回归与分类任务上效果媲美网格搜索,计算成本大幅降低。
  • 揭示不同学界对冷后验效应的关注差异,指导实际应用选择。

冷后验效应(CPE)是贝叶斯深度学习中的一个现象,即降低后验分布的温度通常能提升后验预测分布(PPD)的预测性能。尽管术语暗示更冷的温度更好,但学界逐渐认识到这并非普遍成立。目前仍缺乏系统性方法来确定最优温度,仅依赖网格搜索。本文提出一种数据驱动方法,将温度作为模型参数,直接从数据中估计使测试对数预测密度最大化的最优温度。实验表明,该方法在回归和分类任务上表现接近网格搜索,但计算成本仅为后者的极小部分。最后,我们指出贝叶斯深度学习与广义贝叶斯两个领域对CPE的看法存在分歧:前者关注PPD的预测性能,后者重视模型误设下后验的实用性;两者目标不同,导致对温度的选择偏好各异。

原文摘要 · Abstract (English)

The Cold Posterior Effect (CPE) is a phenomenon in Bayesian Deep Learning (BDL), where tempering the posterior to a cold temperature often improves the predictive performance of the posterior predictive distribution (PPD). Although the term `CPE' suggests colder temperatures are inherently better, the BDL community increasingly recognizes that this is not always the case. Despite this, there remains no systematic method for finding the optimal temperature beyond grid search. In this work, we propose a data-driven approach to select the temperature that maximizes test log-predictive density, treating the temperature as a model parameter and estimating it directly from the data. We empirically demonstrate that our method performs comparably to grid search, at a fraction of the cost, across both regression and classification tasks. Finally, we highlight the differing perspectives on CPE between the BDL and Generalized Bayes communities: while the former primarily emphasizes the predictive performance of the PPD, the latter prioritizes the utility of the posterior under model misspecification; these distinct objectives lead to different temperature preferences.

贝叶斯深度学习温度优化后验估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。