解析高维PLS的理论边界,揭示其在数据融合中的优劣
High-Dimensional Partial Least Squares: Spectral Analysis and Fundamental Limitations
- 基于随机矩阵理论分析跨数据协方差矩阵的奇异向量
- 发现PLS-SVD在特定条件下会失效或表现反常
- 证明其相比独立PCA更优,适用于共性结构提取
偏最小二乘法(PLS)广泛用于高维数据集成,旨在提取跨配对数据集共享的潜在成分。尽管实践成功多年,其在高维场景下的理论理解仍不充分。本文研究一种双高维数据矩阵共享低秩公共潜结构且含个体特异性成分的模型,利用随机矩阵理论分析相关交叉协方差矩阵的奇异向量,推导出估计与真实潜方向对齐的渐近刻画。结果定量解释了基于奇异值分解的PLS(PLS-SVD)的重构性能,并识别出方法表现出反直觉或局限性的区域。基于此分析,对比了PLS-SVD与分别对每组数据应用主成分分析(PCA)的方法,证明其在检测公共潜子空间上具有渐近优势。总体而言,本研究为高维PLS-SVD提供了全面的理论框架,阐明了其优势与根本局限。
原文摘要 · Abstract (English)
Partial Least Squares (PLS) is a widely used method for data integration, designed to extract latent components shared across paired high-dimensional datasets. Despite decades of practical success, a precise theoretical understanding of its behavior in high-dimensional regimes remains limited. In this paper, we study a data integration model in which two high-dimensional data matrices share a low-rank common latent structure while also containing individual-specific components. We analyze the singular vectors of the associated cross-covariance matrix using tools from random matrix theory and derive asymptotic characterizations of the alignment between estimated and true latent directions. These results provide a quantitative explanation of the reconstruction performance of the PLS variant based on Singular Value Decomposition (PLS-SVD) and identify regimes where the method exhibits counter-intuitive or limiting behavior. Building on this analysis, we compare PLS-SVD with principal component analysis applied separately to each dataset and show its asymptotic superiority in detecting the common latent subspace. Overall, our results offer a comprehensive theoretical understanding of high-dimensional PLS-SVD, clarifying both its advantages and fundamental limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。