证明了非线性CCA在特定分布假设下可恢复真实潜在因子。
Provable Affine Identifiability of Nonlinear CCA under Latent Distributional Priors
- 通过源空间分析,将经典统计理论扩展至表示学习。
- 揭示白化是保证映射稳定性的必要条件。
- 解释了近期无对比学习方法的成功原理,适合理论研究者。
本文建立了非线性典型相关分析(CCA)在特定分布先验下可恢复真实潜在因子的充分条件。通过将分析从观测空间转移到源空间,我们把双变量分布的正交多项式展开经典结果推广到表示学习领域,证明了在特定分布假设下具有仿射可识别性。我们严格证明白化是确保学习映射有界且良态的必要条件。此外,我们通过理论证明了岭正则化经验CCA在有限样本条件下收敛于其总体版本。我们的发现为近期基于相关性的无对比学习方法提供了严格的理论基础。在合成数据与渲染图像数据集上的实验及系统消融验证了预测的恢复行为,并展示了违反假设时的失效模式。
原文摘要 · Abstract (English)
In this work, we establish the sufficient conditions under which nonlinear Canonical Correlation Analysis (CCA) recovers ground-truth latent factors up to an affine transformation. By transporting the analysis from the observation space to the source space, we extend classical statistical results on orthogonal polynomial expansions of bivariate distributions to representation learning, proving affine identifiability under specific distributional priors. We formally demonstrate that whitening is strictly necessary to ensure the boundedness and well-conditioning of the learned mappings. Furthermore, we bridge the gap between theory and practice by proving that ridge-regularized empirical CCA converges to its population counterpart in the finite-sample regime. Finally, our findings provide a rigorous theoretical foundation explaining the empirical success of recent correlation-based non-contrastive learning methods. Experiments on synthetic and rendered image datasets, alongside systematic ablations, validate the predicted recovery behavior and illustrate the failure modes that arise when the assumptions are violated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。