arXiv:2604.12340stat.MLcond-mat.stat-mech2026-04

提出无监督学习泛化误差的几何分解方法,揭示误差来源与最优模型结构的关系。

Information-Geometric Decomposition of Generalization Error in Unsupervised Learning

  • 基于信息几何将泛化误差拆解为模型误差、数据偏差和方差三部分。
  • 在ε-PCA中发现最优降维秩等于噪声阈值ε,实现误差最小化。
  • 适用于高维数据建模,尤其适合分析降维算法的泛化性能边界。

我们将无监督学习中的Kullback-Leibler泛化误差(从数据分布到训练模型的期望KL散度)精确分解为三个非负分量:模型误差、数据偏差和方差。该分解对任意e-flat模型类成立,源于信息几何的两个恒等式:广义勾股定理与双重e-混合方差恒等式。以ε-PCA为例进行解析演示——一种正则化主成分分析,其中经验协方差在秩N_K处截断,丢弃方向被固定在噪声底限ε。尽管ε-PCA本身非e-flat,但在各向同性高斯数据下可通过技术重构,保持总泛化误差不变,且各分量具闭式表达。最优秩由λ_{cut}^* = ε确定,即仅保留超过噪声底限的特征值,其选择反映模型误差增益与数据偏差代价之间的边际速率平衡。进一步的边界分析揭示三相图:全保留、中间、坍缩,由下Marchenko-Pastur边缘和可解析计算的坍缩阈值ε_*(α)分隔,其中α为维度与样本数之比。所有结论均经数值验证。

原文摘要 · Abstract (English)

We decompose the Kullback--Leibler generalization error (GE) -- the expected KL divergence from the data distribution to the trained model -- of unsupervised learning into three non-negative components: model error, data bias, and variance. The decomposition is exact for any e-flat model class and follows from two identities of information geometry: the generalized Pythagorean theorem and a dual e-mixture variance identity. As an analytically tractable demonstration, we apply the framework to $ε$-PCA, a regularized principal component analysis in which the empirical covariance is truncated at rank $N_K$ and discarded directions are pinned at a fixed noise floor $ε$. Although rank-constrained $ε$-PCA is not itself e-flat, it admits a technical reformulation with the same total GE on isotropic Gaussian data, under which each component of the decomposition takes closed form. The optimal rank emerges as the cutoff $λ_{\mathrm{cut}}^{*} = ε$ -- the model retains exactly those empirical eigenvalues exceeding the noise floor -- with the cutoff reflecting a marginal-rate balance between model-error gain and data-bias cost. A boundary comparison further yields a three-regime phase diagram -- retain-all, interior, and collapse -- separated by the lower Marchenko--Pastur edge and an analytically computable collapse threshold $ε_{*}(α)$, where $α$ is the dimension-to-sample-size ratio. All claims are verified numerically.

泛化误差信息几何降维ε-PCA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。