arXiv:2505.24781stat.MLcs.CE2025-05

提出高效估算正则化泰勒M估计器的收缩系数方法,显著提升计算速度与精度。

Efficient Estimation of Regularized Tyler's M-Estimator Using Approximate LOOCV

  • 用近似留一法交叉验证优化收缩系数,避免重复运行估计过程
  • 计算复杂度从O(n²)降至O(n),大幅提速且保持高精度
  • 适用于高维重尾数据,如人脸识别、手写数字识别等场景

本文研究正则化泰勒M估计器(RTME)中正则化参数α∈(0,1)的高效估计问题。提出将α设为留一法交叉验证(LOOCV)对数似然损失的最优解,以获得最佳收缩系数。由于传统LOOCV在中等样本量n下计算成本过高,本文提出一种计算高效的近似方法,无需对每个被移除样本重复执行RTME过程。该近似使LOOCV的运行时间复杂度降低至O(n),实现显著加速。实验在重尾椭球分布生成的合成高维数据及真实高维数据集(物体识别、人脸识别、手写数字识别)上验证了方法的有效性,结果表明该方法在效率和准确性上均优于现有文献中的其他方法。

原文摘要 · Abstract (English)

We consider the problem of estimating a regularization parameter, or a shrinkage coefficient $α\in (0,1)$ for Regularized Tyler's M-estimator (RTME). In particular, we propose to estimate an optimal shrinkage coefficient by setting $α$ as the solution to a suitably chosen objective function; namely the leave-one-out cross-validated (LOOCV) log-likelihood loss. Since LOOCV is computationally prohibitive even for moderate sample size $n$, we propose a computationally efficient approximation for the LOOCV log-likelihood loss that eliminates the need for invoking the RTME procedure $n$ times for each sample left out during the LOOCV procedure. This approximation yields an $O(n)$ reduction in the running time complexity for the LOOCV procedure, which results in a significant speedup for computing the LOOCV estimate. We demonstrate the efficiency and accuracy of the proposed approach on synthetic high-dimensional data sampled from heavy-tailed elliptical distributions, as well as on real high-dimensional datasets for object recognition, face recognition, and handwritten digit's recognition. Our experiments show that the proposed approach is efficient and consistently more accurate than other methods in the literature for shrinkage coefficient estimation.

统计估计高维数据正则化交叉验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。