arXiv:2510.11546stat.MLcs.LG2025-10

针对高维回归中的异常值问题,提出一种无需调参的鲁棒估计方法。

Efficient Group Lasso Regularized Rank Regression with Simulation-Based Tuning

  • 用秩损失+组稀疏正则提升抗异常值能力
  • 在小样本下仍保证误差有界,优于主流方法
  • 算法高效可扩展,适合真实大数据场景

高维回归常受重尾噪声和异常值影响,严重削弱最小二乘法的可靠性。为此,采用非光滑的Wilcoxon秩损失函数并引入组稀疏正则化。通过拓展原本仅适用于秩Lasso的免调参性质,提出基于模拟的调参规则,并建立了所提估计量的有限样本误差界。为求解相关优化问题,开发了一种近端增广拉格朗日法,通过证明非多面体KKT映射的度量次正则性,提供新颖的收敛性分析,同时支持子问题的高效半光滑牛顿更新。大量数值实验表明,该估计器在对抗多种领先方法时表现出更强的鲁棒性与有效性,并在模拟与真实数据设置中展现出优于现有最优基线的效率与可扩展性。

原文摘要 · Abstract (English)

High-dimensional regression often suffers from heavy-tailed noise and outliers, which can severely undermine the reliability of least-squares based methods. To improve robustness, we adopt a non-smooth Wilcoxon score based rank objective and incorporate the group sparsity regularization. By extending the tuning-free property originally developed for the rank Lasso, we introduce a simulation-based tuning rule and further establish a finite-sample error bound for the resulting estimator. To solve the associated optimization problem, we develop a proximal augmented Lagrangian method, for which we provide a novel convergence analysis by proving the metric subregularity of the underlying non-polyhedral KKT mapping, while enabling efficient semismooth Newton updates for the subproblems. Extensive numerical experiments demonstrate the robustness and effectiveness of our proposed estimator against several leading alternatives, and showcase the efficiency and scalability of our algorithm compared to the state-of-the-art baseline in both simulated and real-data settings.

高维回归鲁棒估计组稀疏优化算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。