arXiv:2507.16682stat.MLcs.LG2025-07

提出SEDA算法,通过调整数据谱结构提升高维LDA分类效果

Structural Effect and Spectral Enhancement of High-Dimensional Regularized Linear Discriminant Analysis

  • 基于非渐近误分率逼近,揭示数据结构对分类性能的影响
  • 新算法优化协方差矩阵的尖峰特征值,显著降低误分率
  • 适合高维数据分类与降维,尤其在小样本场景表现突出

正则化线性判别分析(RLDA)广泛用于分类与降维,但在高维场景下性能不稳定。现有理论分析难以揭示数据结构对分类效果的影响。本文推导了误分率的非渐近近似,进而分析了结构效应与调整策略。基于此,提出谱增强判别分析(SEDA)算法,通过调整总体协方差矩阵的尖峰特征值来优化数据结构。借助随机矩阵理论中新获得的特征向量结果,推导出SEDA误分率的渐近近似,并得到偏差校正算法与参数选择策略。在合成与真实数据集上的实验表明,SEDA在分类准确率和降维效果上均优于现有LDA方法。

原文摘要 · Abstract (English)

Regularized linear discriminant analysis (RLDA) is a widely used tool for classification and dimensionality reduction, but its performance in high-dimensional scenarios is inconsistent. Existing theoretical analyses of RLDA often lack clear insight into how data structure affects classification performance. To address this issue, we derive a non-asymptotic approximation of the misclassification rate and thus analyze the structural effect and structural adjustment strategies of RLDA. Based on this, we propose the Spectral Enhanced Discriminant Analysis (SEDA) algorithm, which optimizes the data structure by adjusting the spiked eigenvalues of the population covariance matrix. By developing a new theoretical result on eigenvectors in random matrix theory, we derive an asymptotic approximation on the misclassification rate of SEDA. The bias correction algorithm and parameter selection strategy are then obtained. Experiments on synthetic and real datasets show that SEDA achieves higher classification accuracy and dimensionality reduction compared to existing LDA methods.

降维判别分析高维数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。