提出流式PCA的精准收敛与不确定性量化方法,适用于一般秩情形。
Inference and Uncertainty Quantification for Streaming $r$-PCA
- 通过线性化分析,消除旧有理论中的余项问题,获得尖锐收敛率。
- 在稠尾模型下达到极小最大误差率,且在稀疏与稠密尾部均成立。
- 开发在线乘子自助法,适用于高维流式数据的统计推断。
针对流式主成分分析(PCA)中基于Oja算法的两个开放问题:一般秩情形下的尖锐算子范数收敛性,以及子空间估计量的分布推断。现有分析在单秩情况下仍需有界数据假设或存在不衰减的余项,无法适应多项式衰减的谱尾;而分布结果仅限于单秩情形。本文的收敛理论消除了这些余项,得到尖锐速率,在稠尾带状协方差模型下匹配极小最大率(对数因子内)。更一般地,我们在弱非退化条件下证明了跨稠尾与稀尾情形的匹配下界(对数因子内)。分析导出Oja迭代的线性化形式,进而实现高维子空间估计误差的高斯近似,并给出显式极限协方差。还建立了对齐差分在凸集上的逐行高斯近似,恢复了先前单秩结果。为实际推断,我们设计了在线乘子自助算法并证明其一致性。该技术亦可推广至非凸随机逼近的高斯近似与自助推断。
原文摘要 · Abstract (English)
We address two open questions in streaming PCA via Oja's algorithm: sharp operator-norm convergence for general rank under sub-Gaussian data, and distributional inference for the resulting subspace estimator. Existing convergence analyses, even in the rank-one case, either assume bounded data or leave non-vanishing remainder terms that prevent adaptation to a polynomially vanishing tail spectrum, while existing distributional results are confined to the rank-one case. Our convergence theory removes these remainder terms and yields a sharp rate. In the dense-tail spiked covariance regime, this rate matches the minimax rate up to logarithmic factors. More generally, we prove a matching lower bound, up to logarithmic factors, across both dense-tail and sparse-tail regimes under a mild nondegeneracy condition. The analysis yields a linearization of Oja's iterates, which in turn enables a high-dimensional Gaussian approximation for the general-rank subspace estimation error with an explicit limiting covariance. We also establish a row-wise Gaussian approximation over convex sets for the aligned difference, recovering prior rank-one results as special cases. For practical inference, we develop an online multiplier bootstrap algorithm and prove its consistency. Beyond streaming PCA, our techniques contribute to Gaussian approximation and bootstrap inference for nonconvex stochastic approximation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。