让高维输出的高斯过程模型更高效,同时保留输出间依赖关系。
Transformed Latent Variable Multi-Output Gaussian Processes

- 用可学习的嵌入网络映射输入和输出潜变量,构建灵活深度核函数。
- 在超过1万维输出的气候数据上预测准确率超越基线,计算更快。
- 适合需要建模大量相关输出的任务,如多变量时空建模。
多输出高斯过程(MOGP)为建模相关输出提供了一个严谨的概率框架,但在高维输出空间的数据集上面临可扩展性瓶颈。现有方法通常依赖低秩或可分离核等限制性假设,影响表达能力。本文提出变换潜变量多输出高斯过程(T-LVMOGP),可在保持输出间有意义依赖关系的前提下,将MOGP扩展至大规模输出场景。T-LVMOGP通过一个利普希茨正则化神经网络,将输入与输出特异性潜变量映射到嵌入空间,构建灵活的多输出深度核。结合随机变分推断,该模型能有效处理高维输出设置。在多个基准测试中,包括包含超过10,000个输出的气候建模和零膨胀空间转录组数据,T-LVMOGP在预测精度和计算效率方面均优于基线方法。
原文摘要 · Abstract (English)
Multi-Output Gaussian Processes (MOGPs) provide a principled probabilistic framework for modelling correlated outputs but face scalability bottlenecks when applied to datasets with high-dimensional output spaces. To maintain tractability, existing methods typically resort to restrictive assumptions, such as employing low-rank or sum-of-separable kernels, which can limit expressiveness. We propose the Transformed Latent Variable MOGP (T-LVMOGP), a novel framework that scales MOGPs to a massive number of outputs while preserving the capacity to capture meaningful inter-output dependencies. T-LVMOGP constructs a flexible multi-output deep kernel by mapping inputs and output-specific latent variables into an embedding space using a Lipschitz-regularised neural network. Combined with stochastic variational inference, our model effectively scales to high-dimensional output settings. Across diverse benchmarks, including climate modelling with over 10,000 outputs and zero-inflated spatial transcriptomics data, T-LVMOGP outperforms baselines in both predictive accuracy and computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。