让高斯过程在千万级数据上精确计算,不依赖近似方法
gp2Scale: A Class of Compactly Supported Non-Stationary Kernels and Distributed Computing for Exact Gaussian Processes on 10 Million Data Points
- 用可定制的非平稳核发现协方差矩阵天然稀疏结构
- 在1000万数据上实现精确训练,无需诱导点或插值近似
- 适合需要灵活核函数设计的现代高斯过程应用
尽管已有大量工作致力于提升高斯过程的扩展性,但计算速度、预测精度与不确定性量化准确性之间的权衡依然存在。现有方法普遍依赖各类近似,降低精度并限制核函数和噪声模型的设计灵活性——这在表达性强的非平稳核日益流行的当下已成为不可接受的缺陷。本文提出名为gp2Scale的方法,可在不使用诱导点、核插值或邻域近似的情况下,将精确高斯过程扩展至超过1000万数据点,转而利用高斯过程本身的核设计能力。通过高度灵活、紧支撑且非平稳的核函数,自然产生协方差矩阵的稀疏结构,进而用于线性系统求解与对数行列式计算。我们在多个真实数据集上验证了该方法的有效性,并与当前最先进的近似算法进行对比。结果显示,该方法在多数情况下具有更优的近似性能,其真正优势在于对任意高斯过程自定义(核心核函数、噪声模型、均值函数及输入空间类型)的无偏性,使其最适配现代高斯过程应用。
原文摘要 · Abstract (English)
Despite a large corpus of recent work on scaling up Gaussian processes, a stubborn trade-off between computational speed, prediction and uncertainty quantification accuracy, and customizability persists. This is because the vast majority of existing methodologies exploit various levels of approximations that lower accuracy and limit the flexibility of kernel and noise-model designs -- an unacceptable drawback at a time when expressive non-stationary kernels are on the rise in many fields. Here, we propose a methodology we term \emph{gp2Scale} that scales exact Gaussian processes to more than 10 million data points without relying on inducing points, kernel interpolation, or neighborhood-based approximations, and instead leveraging the existing capabilities of a GP: its kernel design. Highly flexible, compactly supported, and non-stationary kernels lead to the identification of naturally occurring sparse structure in the covariance matrix, which is then exploited for the calculations of the linear system solution and the log-determinant for training. We demonstrate our method's functionality on several real-world datasets and compare it with state-of-the-art approximation algorithms. Although we show superior approximation performance in many cases, the method's real power lies in its agnosticism toward arbitrary GP customizations -- core kernel design, noise, and mean functions -- and the type of input space, making it optimally suited for modern Gaussian process applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。