arXiv:2505.13723cs.LGmath.OC2025-05NeurIPS被引 7

提出ADASAP算法,让高斯过程在超大规模数据上也能快速推理。

Turbocharging Gaussian Process Inference with Approximate Sketch-and-Project

  • 用近似压缩投影法加速求解高斯过程的线性系统
  • 在超过3亿样本的数据集上实现高效推理,突破现有方法极限
  • 适合需要大规模贝叶斯推断的研究者,如科学机器学习与优化

高斯过程(GPs)在生物统计、科学机器学习和贝叶斯优化中至关重要,因其能提供概率预测和不确定性建模。然而,其推断难以扩展到大规模数据,因需求解规模随样本数平方增长的线性系统。本文提出一种近似、分布式、加速的压缩投影算法(ADASAP),显著提升可扩展性。基于行列式点过程理论,证明了压缩投影法诱导的后验均值可快速收敛至真实后验均值,首次实现对顶部谱基函数的后验均值估计且无需依赖条件数,表明方法在高斯过程推断中的合理性。ADASAP在多个基准数据集及大规模贝叶斯优化任务中超越共轭梯度与坐标下降等先进求解器,成功处理超过3×10⁸样本的数据集,为该领域首次达成。

原文摘要 · Abstract (English)

Gaussian processes (GPs) play an essential role in biostatistics, scientific machine learning, and Bayesian optimization for their ability to provide probabilistic predictions and model uncertainty. However, GP inference struggles to scale to large datasets (which are common in modern applications), since it requires the solution of a linear system whose size scales quadratically with the number of samples in the dataset. We propose an approximate, distributed, accelerated sketch-and-project algorithm ($\texttt{ADASAP}$) for solving these linear systems, which improves scalability. We use the theory of determinantal point processes to show that the posterior mean induced by sketch-and-project rapidly converges to the true posterior mean. In particular, this yields the first efficient, condition number-free algorithm for estimating the posterior mean along the top spectral basis functions, showing that our approach is principled for GP inference. $\texttt{ADASAP}$ outperforms state-of-the-art solvers based on conjugate gradient and coordinate descent across several benchmark datasets and a large-scale Bayesian optimization task. Moreover, $\texttt{ADASAP}$ scales to a dataset with $> 3 \cdot 10^8$ samples, a feat which has not been accomplished in the literature.

高斯过程大规模推理加速算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。