arXiv:2505.07272stat.MLcs.LG2025-05被引 7

ALPCAH可自适应估计每样本噪声,提升低秩子空间建模精度。

ALPCAH: Subspace Learning for Sample-wise Heteroscedastic Data

  • 通过自适应估计每样本噪声方差,改进子空间基学习
  • 无需先验噪声信息或子空间维度,支持软秩约束
  • 适用于噪声质量差异大的真实数据,尤其适合高维异质数据

主成分分析(PCA)是数据降维的重要工具。然而,某些应用中的数据因样本间噪声特性不同而呈现异质性。现有异方差方法旨在处理此类数据质量不一的问题。本文提出一种名为ALPCAH的子空间学习方法,能够估计每个样本的噪声方差,并利用该信息优化数据低秩结构对应的子空间基估计。本方法不假设低秩成分的分布形式,也不依赖已知的噪声方差。此外,采用软秩约束,无需事先知道子空间维度。为进一步提升效率,本文还提出矩阵分解版本的LR-ALPCAH,速度更快、内存占用更低,但需预先知晓或估计子空间维度。模拟实验与真实数据测试表明,考虑数据异方差性显著优于现有算法。代码已公开于https://github.com/javiersc1/ALPCAH。

原文摘要 · Abstract (English)

Principal component analysis (PCA) is a key tool in the field of data dimensionality reduction. However, some applications involve heterogeneous data that vary in quality due to noise characteristics associated with each data sample. Heteroscedastic methods aim to deal with such mixed data quality. This paper develops a subspace learning method, named ALPCAH, that can estimate the sample-wise noise variances and use this information to improve the estimate of the subspace basis associated with the low-rank structure of the data. Our method makes no distributional assumptions of the low-rank component and does not assume that the noise variances are known. Further, this method uses a soft rank constraint that does not require subspace dimension to be known. Additionally, this paper develops a matrix factorized version of ALPCAH, named LR-ALPCAH, that is much faster and more memory efficient at the cost of requiring subspace dimension to be known or estimated. Simulations and real data experiments show the effectiveness of accounting for data heteroscedasticity compared to existing algorithms. Code available at https://github.com/javiersc1/ALPCAH.

子空间学习异方差降维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。