arXiv:2506.20425stat.MLcs.LG2025-06被引 1

提出可处理上千变量的线性混合模型稀疏选择方法,速度快且有理论保障。

Scalable Subset Selection in Linear Mixed Models

  • 基于ℓ₀正则化与坐标下降,实现高效变量筛选
  • 支持数千变量在秒级完成计算,实测表现优异
  • 适合高维异质数据建模,如个性化医疗研究

线性混合模型(LMM)通过固定效应和随机效应联合分析异质数据,在个性化医疗等领域至关重要。当前数据规模日益庞大,常含数千个候选预测变量,亟需稀疏化以提升预测与可解释性。然而,现有针对LMM的稀疏学习方法难以扩展至数百以上变量,远落后于忽略随机效应的线性模型方法。本文提出一种新的ℓ₀正则化方法,可在数秒至数分钟内完成含数千变量的LMM子集选择。计算方面,设计坐标下降算法并提供收敛性保证;引入局部搜索算法以应对非凸优化问题。两者均可通过惩罚似然近似拓展至广义LMM。统计方面,给出该方法在有限样本下对Kullback-Leibler散度的界。实验表明其在合成与真实数据上均表现优秀。

原文摘要 · Abstract (English)

Linear mixed models (LMMs), which incorporate fixed and random effects, are key tools for analyzing heterogeneous data, such as in personalized medicine. Nowadays, this type of data is increasingly wide, sometimes containing thousands of candidate predictors, necessitating sparsity for prediction and interpretation. However, existing sparse learning methods for LMMs do not scale well beyond tens or hundreds of predictors, leaving a large gap compared with sparse methods for linear models, which ignore random effects. This paper closes the gap with a new $\ell_0$ regularized method for LMM subset selection that can run on datasets containing thousands of predictors in seconds to minutes. On the computational front, we develop a coordinate descent algorithm as our main workhorse and provide a guarantee of its convergence. We also develop a local search algorithm to help traverse the nonconvex optimization surface. Both algorithms readily extend to subset selection in generalized LMMs via a penalized quasi-likelihood approximation. On the statistical front, we provide a finite-sample bound on the Kullback-Leibler divergence of the new method. We then demonstrate its excellent performance in experiments involving synthetic and real datasets.

线性混合模型稀疏学习变量选择高维数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。