针对高维强噪声数据,提出两种自适应采样方法,显著提升回归稳定性与精度。
Adaptive and Stratified Subsampling for High-Dimensional Robust Estimation
- 采用自适应重要性采样与分层采样策略,应对重尾噪声和数据依赖问题。
- 在20%污染下误差比均匀采样低3.1倍,Riboflavin数据集测试MSE降低29.5%。
- 适用于高维金融、基因组等存在异常值和时间依赖的数据分析场景。
我们研究在有限方差重尾噪声、epsilon-污染及alpha-混合依赖条件下,基于自适应重要性采样(AIS)和分层采样(SS)的鲁棒高维稀疏回归。在精确界定的子高斯设计与有限方差噪声下,采样规模m可达到极小最大最优率。理论闭环:定理4.6适用于终止时权重稳定的AIS(命题4.1),而SS符合Lecue与Lerasle的中位数均值M-估计框架(命题4.3)。通过新的稀疏精度假设下的节点Lasso精度估计器完成去偏化,实现坐标层面的有效置信区间(定理4.14)。alpha-混合扩展采用日历块协议以保障时间分离(定理4.12)。实验表明,在20%污染下,AIS误差仅为均匀采样的3.10倍;在Riboflavin数据集(p=4,088,n=71)上测试MSE降低29.5%。
原文摘要 · Abstract (English)
We study robust high-dimensional sparse regression under finite-variance heavy-tailed noise, epsilon-contamination, and alpha-mixing dependence via two subsampling estimators: Adaptive Importance Sampling (AIS) and Stratified Sub-sampling (SS). Under sub-Gaussian design whose scopeis precisely delimited and finite-variance noise, a subsample of size m achieves the minimax-optimal rate. We close the theory-algorithm gap: Theorem 4.6 applies to AIS at termination conditional on stabilized weights (Proposition 4.1), and SS fits the median-of-means M-estimation framework of Lecue and Lerasle (Proposition 4.3). The de-biasing step is fully specified via the nodewise-Lasso precision estimator under a new sparse-precision assumption, yielding valid coordinate-wise CIs (Theorem 4.14). The alpha-mixing extension uses a calendar-time block protocol that guarantees temporal separation (Theorem 4.12). Empirically, AIS achieves 3.10 times lower error than uniform subsampling at 20% contamination, and 29.5% lower test MSE on Riboflavin (p=4,088 and n=71).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。