arXiv:2507.05064stat.MLcs.LG2025-07被引 2

提出VIF方法,让高斯过程高效处理大规模数据。

Vecchia-Inducing-Points Full-Scale Approximations for Gaussian Processes

  • 结合局部邻近与全局诱导点,用相关性找邻居优化计算
  • 在真实和模拟数据上比现有方法更快更准更稳定
  • 适合需要高效高斯过程建模的研究者使用

高斯过程是灵活的非参数概率模型,但其在大规模数据上的可扩展性受计算限制。为此,我们提出结合全局诱导点与局部Vecchia近似的全尺度近似方法(VIF)。该方法利用基于相关性的邻居查找策略,通过改进的覆盖树算法实现对残差过程的高效近似。进一步地,针对非高斯似然,引入迭代方法,在拉普拉斯近似下使训练与预测的计算成本相比基于Cholesky分解的方法降低多个数量级。我们提出了新型预条件器并给出理论收敛结果。在模拟与真实数据集上的大量实验表明,VIF近似在计算效率、精度与数值稳定性方面均优于当前最优方法。所有方法已集成于开源C++库GPBoost,并提供高级别Python与R接口。

原文摘要 · Abstract (English)

Gaussian processes are flexible, probabilistic, non-parametric models widely used in machine learning and statistics. However, their scalability to large data sets is limited by computational constraints. To overcome these challenges, we propose Vecchia-inducing-points full-scale (VIF) approximations combining the strengths of global inducing points and local Vecchia approximations. Vecchia approximations excel in settings with low-dimensional inputs and moderately smooth covariance functions, while inducing point methods are better suited to high-dimensional inputs and smoother covariance functions. Our VIF approach bridges these two regimes by using an efficient correlation-based neighbor-finding strategy for the Vecchia approximation of the residual process, implemented via a modified cover tree algorithm. We further extend our framework to non-Gaussian likelihoods by introducing iterative methods that substantially reduce computational costs for training and prediction by several orders of magnitudes compared to Cholesky-based computations when using a Laplace approximation. In particular, we propose and compare novel preconditioners and provide theoretical convergence results. Extensive numerical experiments on simulated and real-world data sets show that VIF approximations are both computationally efficient as well as more accurate and numerically stable than state-of-the-art alternatives. All methods are implemented in the open source C++ library GPBoost with high-level Python and R interfaces.

高斯过程大规模推断近似方法机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。