arXiv:2602.10680stat.MLcond-mat.dis-nn2026-02被引 6

非线性自编码器能发现主成分分析遗漏的隐藏结构。

A solvable high-dimensional model where nonlinear autoencoders learn structure invisible to PCA while test loss misaligns with generalization

  • 设计可解析的高维模型,区分可见与隐藏的潜在因子
  • 非线性自编码器准确恢复线性方法无法捕捉的结构
  • 揭示测试损失与泛化能力可能不一致的机制

真实数据常包含无法通过特征间线性相关性检测的隐藏结构。例如,潜在因素可能以协调方式影响数据,但其效应对基于协方差的方法(如主成分分析,PCA)不可见。实践中,非线性神经网络在无监督和自监督学习中常能提取此类结构。然而,构建一个可严格分析该优势的最小高维模型仍是一个开放的理论挑战。本文提出一个可解析的高维尖峰模型,包含两个潜在因子:一个可通过协方差检测,另一个统计相关但不相关,仅体现在高阶矩中。PCA 和线性自编码器无法恢复后者,而一个最小非线性自编码器可证明地同时提取两者。我们分析了总体风险和经验风险最小化。该模型还提供了一个可解析示例,说明自监督测试损失与表示质量严重错位:非线性自编码器恢复了线性方法遗漏的潜在结构,尽管其重建损失更高。

原文摘要 · Abstract (English)

Many real-world datasets contain hidden structure that cannot be detected by simple linear correlations between input features. For example, latent factors may influence the data in a coordinated way, even though their effect is invisible to covariance-based methods such as PCA. In practice, nonlinear neural networks often succeed in extracting such hidden structure in unsupervised and self-supervised learning. However, constructing a minimal high-dimensional model where this advantage can be rigorously analyzed has remained an open theoretical challenge. We introduce a tractable high-dimensional spiked model with two latent factors: one visible to covariance, and one statistically dependent yet uncorrelated, appearing only in higher-order moments. PCA and linear autoencoders fail to recover the latter, while a minimal nonlinear autoencoder provably extracts both. We analyze both the population risk, and empirical risk minimization. Our model also provides a tractable example where self-supervised test loss is poorly aligned with representation quality: nonlinear autoencoders recover latent structure that linear methods miss, even though their reconstruction loss is higher.

非线性降维自编码器高维统计表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。