arXiv:2603.13806stat.MLcs.LG2026-03被引 2

提出自适应稀疏主成分分析法,提升高维数据建模的稳定性与可解释性

An Interpretable and Stable Framework for Sparse Principal Component Analysis

  • 引入均衡参数动态调节变量惩罚,实现稀疏性与方差解释力的灵活权衡
  • 在高维噪声数据中显著提升特征选择准确率,保持更高累计方差解释能力
  • 适合需要稳定特征筛选与清晰解释的高维数据分析场景

稀疏主成分分析(SPCA)旨在解决高维数据中主成分分析(PCA)可解释性差和变量冗余的问题。然而,传统SPCA对所有变量施加统一惩罚,未考虑变量重要性差异,导致在高噪声或结构复杂环境下性能不稳定。本文提出SP-SPCA,通过在正则化框架中引入单一均衡参数,自适应调整变量惩罚强度。该方法对L2惩罚进行改造,可在保持计算效率的同时灵活控制稀疏性与解释方差的权衡。模拟研究表明,所提方法在识别稀疏载荷模式、过滤噪声变量及保留累积方差方面均优于标准SPCA,尤其在高维和高噪声条件下表现更优。在犯罪数据与金融市场数据的实际应用中,该方法选出更少但更相关的变量,降低模型复杂度的同时维持强解释力。总体而言,该方法为复杂高维数据提供了鲁棒、高效的稀疏建模方案,在稳定性、特征选择和可解释性上具有明显优势。

原文摘要 · Abstract (English)

Sparse principal component analysis (SPCA) addresses the poor interpretability and variable redundancy often encountered by principal component analysis (PCA) in high-dimensional data. However, SPCA typically imposes uniform penalties on variables and does not account for differences in variable importance, which may lead to unstable performance in highly noisy or structurally complex settings. We propose SP-SPCA, a method that introduces a single equilibrium parameter into the regularization framework to adaptively adjust variable penalties. This modification of the L2 penalty provides flexible control over the trade-off between sparsity and explained variance while maintaining computational efficiency. Simulation studies show that the proposed method consistently outperforms standard sparse principal component methods in identifying sparse loading patterns, filtering noise variables, and preserving cumulative variance, especially in high-dimensional and noisy settings. Empirical applications to crime and financial market data further demonstrate its practical utility. In real data analyses, the method selects fewer but more relevant variables, thereby reducing model complexity while maintaining explanatory power. Overall, the proposed approach offers a robust and efficient alternative for sparse modeling in complex high-dimensional data, with clear advantages in stability, feature selection, and interpretability

稀疏主成分高维数据可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。