arXiv:2410.18973stat.COcs.LG2024-10中稿 · the 41st Conferenc…

无需调参的高效贝叶斯核数据集构造方法

Tuning-Free Coreset Markov Chain Monte Carlo via Hot DoG

  • 提出Hot DoG方法,避免传统优化中的学习率调参问题
  • 在多个数据集上优于其他免调参优化方法,媲美最优ADAM
  • 适合希望省去超参调优、追求稳定性能的贝叶斯推断用户

贝叶斯核数据集(Bayesian coreset)是原始数据集的小型加权子集,可在推断中替代全量数据以降低计算成本。当前最先进的核数据集构建算法Coreset MCMC利用针对核数据集后验的自适应马尔可夫链采样,通过随机梯度优化训练核数据集权重。然而,该方法的核数据集质量及其后验近似精度对随机优化的学习率敏感。本文提出一种免学习率调整的随机梯度优化方法——热启动梯度距离(Hot DoG),用于在Coreset MCMC中训练核数据集权重,无需人工调参。我们提供了Hot DoG生成核数据集权重的收敛性理论分析。实验结果表明,Hot DoG在多个数据集上均优于其他免调参的随机梯度方法,且性能与经过最优调参的ADAM相当。

原文摘要 · Abstract (English)

A Bayesian coreset is a small, weighted subset of a data set that replaces the full data during inference to reduce computational cost. The state-of-the-art coreset construction algorithm, Coreset Markov chain Monte Carlo (Coreset MCMC), uses draws from an adaptive Markov chain targeting the coreset posterior to train the coreset weights via stochastic gradient optimization. However, the quality of the constructed coreset, and thus the quality of its posterior approximation, is sensitive to the stochastic optimization learning rate. In this work, we propose a learning-rate-free stochastic gradient optimization procedure, Hot-start Distance over Gradient (Hot DoG), for training coreset weights in Coreset MCMC without user tuning effort. We provide a theoretical analysis of the convergence of the coreset weights produced by Hot DoG. We also provide empirical results demonstrate that Hot DoG provides higher quality posterior approximations than other learning-rate-free stochastic gradient methods, and performs competitively to optimally-tuned ADAM.

贝叶斯推断核数据集无调参优化马尔可夫链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。