用高斯过程改进密度估计,可闭式求解且适合大规模数据。
Gaussian Process Tilted Nonparametric Density Estimation using Fisher Divergence Score Matching
- 基于高斯过程的指数修正模型,结合随机傅里叶特征实现闭式优化。
- 三种基于费舍尔散度的目标函数均能闭式求解,精度优于基准方法。
- 无需迭代训练,仅需单遍扫描数据,适合大数据场景。
本文提出一种基于高斯过程(GP)的非参数密度估计器,通过将多元正态分布与指数化的GP修正项相乘构建,称为GP-倾斜非参数密度。利用随机傅里叶特征(RFF)近似将GP得分表示为线性形式,推导出三种基于费舍尔散度(FD)得分匹配的闭式学习算法:基础版、噪声条件版及基于变分推断(VI)的新方法。为此,我们提出类似ELBO的优化目标以逼近后验分布,并得到费舍尔变分预测分布。RFF表示等价于带余弦激活的单层神经网络得分模型,使所有期望可解析求解;高斯基底分布则保障了变分推断的可处理性并确保密度定义良好。在多个低维密度估计任务中验证了三种算法及一种最大后验(MAP)基线。由于学习问题具有闭式解,无需依赖迭代算法,仅需单次遍历数据收集充分统计量,特别适用于大规模数据集。
原文摘要 · Abstract (English)
We propose a nonparametric density estimator based on the Gaussian process (GP) and derive three novel closed form learning algorithms based on Fisher divergence (FD) score matching. The density estimator is formed by multiplying a base multivariate normal distribution with an exponentiated GP refinement, and so we refer to it as a GP-tilted nonparametric density. By representing the GP part of the score as a linear function using the random Fourier feature (RFF) approximation, we show that optimization can be solved in closed form for the three FD-based objectives considered. This includes the basic and noise conditional versions of the Fisher divergence, as well as an alternative to noise conditional FD models based on variational inference (VI) that we propose in this paper. For this novel learning approach, we propose an ELBO-like optimization to approximate the posterior distribution, with which we then derive a Fisher variational predictive distribution. The RFF representation of the GP, which is functionally equivalent to a single layer neural network score model with cosine activation, provides a useful linear representation of the GP for which all expectations can be solved. The Gaussian base distribution also helps with tractability of the VI approximation and ensures that our proposed density is well-defined. We demonstrate our three learning algorithms, as well as a MAP baseline algorithm, on several low dimensional density estimation problems. The closed form nature of the learning problem removes the reliance on iterative learning algorithms, making this technique particularly well-suited to big data sets, since only sufficient statistics collected from a single pass through the data is needed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。