提出可同时学习均值与方差的梯度提升混合模型,提升集群数据预测精度。
Gradient Boosted Mixed Models: Flexible Estimation of Mean and Variance Components for Clustered Data
- 用似然梯度联合学习固定效应与随机效应的非参数函数
- 在真实数据上显著降低随机效应方差估计误差,降幅达8倍
- 适合需要个性化预测与置信区间的研究者,如医学临床试验
我们提出一种将梯度提升与混合效应模型结合的新方法——梯度提升混合模型(GBMixed),通过基于似然的梯度联合学习响应变量的均值与方差成分。GBMixed 能够非参数化地估计整体均值的固定效应函数,并灵活地让随机效应协方差矩阵及残差方差依赖于协变量。实验表明,该方法能准确恢复复杂非线性固定效应函数以及协变量相关的协方差结构,在线性混合模型、自然梯度提升和高斯过程提升等多种方法中表现出更优的点预测与概率预测性能。在模拟中,当方差分量随协变量变化时,GBMixed 相比线性混合模型与高斯过程提升,随机效应方差函数的均方误差降低八倍。
原文摘要 · Abstract (English)
We introduce a novel way to combine gradient boosting with mixed effects models, whereby the mean and variance components are learned jointly as functions of covariates via likelihood-based gradients. Gradient Boosted Mixed Models (GBMixed) estimates a nonparametric fixed effects function characterizing the overall mean of the response, while also allowing the random effects covariance matrix along with the residual variance to depend on covariates in a flexible manner. We demonstrate how GBMixed facilitates covariate-dependent random effect predictions, and subsequently point predictions and prediction intervals for individual treatment effects, that can adapt between population-level and cluster-level information. Experiments and applications to two real-world datasets show that GBMixed can accurately recover complex nonlinear fixed effect functions and covariate-dependent covariances in a linear mixed model, while also improving point and probabilistic predictive performance compared with several existing approaches such as parametric linear mixed models, Natural Gradient Boosting, and Gaussian Process Boosting. In simulations where the variance components are designed to vary as a function of covariates, GBMixed reduces the mean squared error in recovering the random effects variance function by a factor of eight relative to linear mixed models and Gaussian Process Boosting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。