研究逻辑回归最大似然估计在有限样本下的表现,揭示其存在性与预测精度的关键影响因素。
Finite-sample performance of the maximum likelihood estimator in logistic regression
- 基于高斯协变量和正确定义模型,给出MLE存在性和风险的精确非渐近界
- 在信号强度与维度条件下,证明了当数据不完全线性可分时MLE可存在且风险可控
- 推广至非高斯、模型误设及伯努利设计场景,揭示参数方向对估计敏感性的关键作用
逻辑回归是描述二元响应与多维协变量概率依赖的经典模型。本文研究逻辑回归最大似然估计(MLE)的预测性能,以逻辑损失风险为评估标准。重点关注两个问题:一是MLE的存在性(当数据不线性可分时存在),二是其存在时的准确性。这些性质取决于协变量维度与信号强度。在高斯协变量且模型正确设定的条件下,我们获得了关于MLE存在性与超出逻辑风险的精确非渐近保证。随后,将结果推广至满足特定二维边缘条件的非高斯协变量,以及可能误设的统计学习一般情形。最后,考虑伯努利设计情况,发现MLE行为对参数方向极为敏感。
原文摘要 · Abstract (English)
Logistic regression is a classical model for describing the probabilistic dependence of binary responses to multivariate covariates. We consider the predictive performance of the maximum likelihood estimator (MLE) for logistic regression, assessed in terms of logistic risk. We consider two questions: first, that of the existence of the MLE (which occurs when the dataset is not linearly separated), and second, that of its accuracy when it exists. These properties depend on both the dimension of covariates and the signal strength. In the case of Gaussian covariates and a well-specified logistic model, we obtain sharp non-asymptotic guarantees for the existence and excess logistic risk of the MLE. We then generalize these results in two ways: first, to non-Gaussian covariates satisfying a certain two-dimensional margin condition, and second to the general case of statistical learning with a possibly misspecified logistic model. Finally, we consider the case of a Bernoulli design, where the behavior of the MLE is highly sensitive to the parameter direction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。