arXiv:2602.02190stat.MLcs.LG2026-02被引 1

研究多概率测度的PCA收敛性,揭示稀疏与密集采样间的转换规律。

PCA of probability measures: Sparse and Dense sampling regimes

  • 在双渐近框架下分析n个测度各由m样本观测时的PCA收敛性
  • 得到误差率n⁻¹/² + m⁻ᵃ,α由嵌入方式决定,揭示采样密度影响
  • 证明密集采样下率可达极小最大最优,适合高维概率数据分析

对概率测度进行主成分分析(PCA)的常见方法是将其嵌入希尔伯特空间,从而应用标准函数型PCA技术。尽管从m个样本估计单个测度的嵌入收敛速率已有充分研究,但关于多个测度的情形尚无系统分析。本文研究在双渐近框架下的PCA:共观测n个概率测度,每个测度通过m个样本获得。我们推导出经验协方差算子和PCA过拟合风险的收敛率形式为n⁻¹/² + m⁻ᵃ,其中α>0依赖于所选嵌入方式。该结果刻画了测度数量n与每测度样本数m之间的关系,揭示了从稀疏(小m)到密集(大m)采样模式的收敛行为转变。此外,我们证明密集情形下的误差率对经验协方差误差达到极小最大最优。数值实验验证了这些理论结果,并表明合理子采样可在降低计算成本的同时保持PCA精度。

原文摘要 · Abstract (English)

A common approach to perform PCA on probability measures is to embed them into a Hilbert space where standard functional PCA techniques apply. While convergence rates for estimating the embedding of a single measure from $m$ samples are well understood, the literature has not addressed the setting involving multiple measures. In this paper, we study PCA in a double asymptotic regime where $n$ probability measures are observed, each through $m$ samples. We derive convergence rates of the form $n^{-1/2} + m^{-α}$ for the empirical covariance operator and the PCA excess risk, where $α>0$ depends on the chosen embedding. This characterizes the relationship between the number $n$ of measures and the number $m$ of samples per measure, revealing a sparse (small $m$) to dense (large $m$) transition in the convergence behavior. Moreover, we prove that the dense-regime rate is minimax optimal for the empirical covariance error. Our numerical experiments validate these theoretical rates and demonstrate that appropriate subsampling preserves PCA accuracy while reducing computational cost.

主成分分析概率测度收敛性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。