用线性混合模型提升变分自编码器对高维纵向数据的建模能力。
Latent mixed-effect models for high-dimensional longitudinal data
- 结合线性混合模型与变分推断,为VAE设计条件先验
- 在模拟与真实数据上表现优于现有方法
- 模型可扩展、可解释,适合有协变量的纵向数据分析
纵向数据建模是重要但具挑战性的任务,常具有高维度、非线性效应和随时间变化的协变量特征。基于高斯过程先验的变分自编码器(VAEs)虽能有效建模时序数据,但训练成本高,且难以充分利用纵向数据中的丰富协变量,限制了实际应用。本文提出LMM-VAE,利用线性混合模型(LMMs)和摊销变分推断构建VAE的条件先验,实现可扩展、可解释且可识别的建模框架。理论分析揭示其与基于高斯过程的方法之间的统一联系。在模拟数据和真实数据集上,该方法性能媲美甚至超越现有技术。
原文摘要 · Abstract (English)
Modelling longitudinal data is an important yet challenging task. These datasets can be high-dimensional, contain non-linear effects and time-varying covariates. Gaussian process (GP) prior-based variational autoencoders (VAEs) have emerged as a promising approach due to their ability to model time-series data. However, they are costly to train and struggle to fully exploit the rich covariates characteristic of longitudinal data, making them difficult for practitioners to use effectively. In this work, we leverage linear mixed models (LMMs) and amortized variational inference to provide conditional priors for VAEs, and propose LMM-VAE, a scalable, interpretable and identifiable model. We highlight theoretical connections between it and GP-based techniques, providing a unified framework for this class of methods. Our proposal performs competitively compared to existing approaches across simulated and real-world datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。