首次为稀疏模型下的惩罚最小截断平方提供有限样本误差界。
Non-asymptotic analysis of the performance of the penalized least trimmed squares in sparse models
- 提出非渐近分析框架,突破传统大样本假设限制。
- 在高维稀疏场景下,给出估计与预测的高概率误差上界。
- 适合关注高维统计推断稳健性的研究者与工程师。
最小截断平方(LTS)估计器是经典最小二乘法的著名鲁棒替代方案,在位置估计、回归、机器学习及人工智能领域广泛应用。尽管已有大量关于LTS的研究,涵盖其鲁棒性、计算算法、非线性扩展及渐近性质等,但其在高维真实数据稀疏模型中的应用中,维度p(达数千)远大于样本量n(数十或数百)。在此类实际场景中,样本量n通常代表具有特定属性的子群体数量(如阿尔茨海默病、帕金森病、白血病或肌萎缩侧索硬化症患者的数量),而总体大小N为有限固定值。因此,假设n趋于无穷的渐近分析在实践中缺乏说服力与合理性。本文首次建立了基于惩罚LTS的有限样本(非渐近)误差界,证明了在高概率下估计与预测的性能保证,为高维稀疏建模提供了更可靠、更实用的理论支持。
原文摘要 · Abstract (English)
The least trimmed squares (LTS) estimator is a renowned robust alternative to the classic least squares estimator and is popular in location, regression, machine learning, and AI literature. Many studies exist on LTS, including its robustness, computation algorithms, extension to non-linear cases, asymptotics, etc. The LTS has been applied in the penalized regression in a high-dimensional real-data sparse-model setting where dimension $p$ (in thousands) is much larger than sample size $n$ (in tens, or hundreds). In such a practical setting, the sample size $n$ often is the count of sub-population that has a special attribute (e.g. the count of patients of Alzheimer's, Parkinson's, Leukemia, or ALS, etc.) among a population with a finite fixed size N. Asymptotic analysis assuming that $n$ tends to infinity is not practically convincing and legitimate in such a scenario. A non-asymptotic or finite sample analysis will be more desirable and feasible. This article establishes some finite sample (non-asymptotic) error bounds for estimating and predicting based on LTS with high probability for the first time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。