arXiv:2410.01093math.STcs.LG2024-10被引 7

高维逻辑回归在缺失数据下,通过正则化与插补实现近似最优预测。

High-dimensional logistic regression with missing data: Imputation, regularization, and universality

  • 基于岭正则化与插补策略处理缺失或噪声协变量。
  • 在独立同分布假设下,精确刻画预测与估计误差,且结果具普适性。
  • 单次插补加正则化可逼近贝叶斯最优性能,适合高维缺失数据场景。

研究高维岭正则化逻辑回归在协变量缺失或受加性噪声污染的情形。当协变量及噪声均独立同正态分布时,我们精确刻画了预测误差与估计误差。此外,证明这些结果具有普适性:只要数据矩阵元素满足独立性和矩条件,结论仍成立。该普适性使得我们能深入分析完全随机缺失情形下的多种插补策略。通过与统计物理中副本理论所预言的贝叶斯最优性能对比,发现:(i) 单次插补与简单多重插补存在本质区别;(ii) 在单次插补逻辑回归中加入简单岭正则化项,其预测误差几乎等同于贝叶斯最优。研究通过大量数值实验验证。

原文摘要 · Abstract (English)

We study high-dimensional, ridge-regularized logistic regression in a setting in which the covariates may be missing or corrupted by additive noise. When both the covariates and the additive corruptions are independent and normally distributed, we provide exact characterizations of both the prediction error as well as the estimation error. Moreover, we show that these characterizations are universal: as long as the entries of the data matrix satisfy a set of independence and moment conditions, our guarantees continue to hold. Universality, in turn, enables the detailed study of several imputation-based strategies when the covariates are missing completely at random. We ground our study by comparing the performance of these strategies with the conjectured performance -- stemming from replica theory in statistical physics -- of the Bayes optimal procedure. Our analysis yields several insights including: (i) a distinction between single imputation and a simple variant of multiple imputation and (ii) that adding a simple ridge regularization term to single-imputed logistic regression can yield an estimator whose prediction error is nearly indistinguishable from the Bayes optimal prediction error. We supplement our findings with extensive numerical experiments.

逻辑回归缺失数据正则化高维统计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。