用汇总数据训练高斯过程模型,解决隐私与成本难题
Learning from Summarized Data: Gaussian Process Regression with Sample Quasi-Likelihood
- 提出样本似然概念,仅用汇总数据实现高斯过程学习
- 精度受汇总粒度与核函数长度尺度的相对关系影响
- 适合需保护隐私的空间数据分析场景
高斯过程回归是一种强大的贝叶斯非线性回归方法。近期研究已能通过非高斯似然捕获多种观测类型,对空间建模有重要价值。然而,当仅能获取包含代表性特征、统计量和数据点数的汇总数据时,仍面临挑战,这通常源于空间数据的隐私顾虑和管理成本。本文在高斯过程回归框架下,针对仅使用汇总数据的学习与推断问题展开研究。分析了利用代表性特征导致的边缘似然与后验分布近似误差,并引入样本拟似然概念,使仅基于汇总数据的学习与推断成为可能。在满足一定假设的非高斯似然可通过刻画样本拟似然函数的方差函数来表示。理论与实验结果表明,近似性能取决于汇总数据粒度与协方差函数长度尺度的相对关系。在真实数据集上的实验验证了该方法在空间建模中的实用性。
原文摘要 · Abstract (English)
Gaussian process regression is a powerful Bayesian nonlinear regression method. Recent research has enabled the capture of many types of observations using non-Gaussian likelihoods. To deal with various tasks in spatial modeling, we benefit from this development. Difficulties still arise when we can only access summarized data consisting of representative features, summary statistics, and data point counts. Such situations frequently occur primarily due to concerns about confidentiality and management costs associated with spatial data. This study tackles learning and inference using only summarized data within the framework of Gaussian process regression. To address this challenge, we analyze the approximation errors in the marginal likelihood and posterior distribution that arise from utilizing representative features. We also introduce the concept of sample quasi-likelihood, which facilitates learning and inference using only summarized data. Non-Gaussian likelihoods satisfying certain assumptions can be captured by specifying a variance function that characterizes a sample quasi-likelihood function. Theoretical and experimental results demonstrate that the approximation performance is influenced by the granularity of summarized data relative to the length scale of covariance functions. Experiments on a real-world dataset highlight the practicality of our method for spatial modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。