arXiv:2601.22756cs.LG2026-01

从嵌入维度和分布收敛看模型泛化,不依赖参数量

Understanding Generalization from Embedding Dimension and Distributional Convergence

  • 用嵌入分布的内在维度和映射敏感度分析泛化性能
  • 嵌入维度越低,预测误差越小,与实证结果一致
  • 适用于评估模型泛化能力,尤其适合跨架构比较

深度神经网络在严重过参数化下仍具有良好的泛化能力,这挑战了传统的基于参数数量的分析方法。本文从表示视角出发,研究学习到的嵌入几何如何控制固定模型的预测性能。我们证明,总体风险可被两个因素界定:(i) 嵌入分布的内在维度,决定经验嵌入分布向总体分布收敛的速度(以Wasserstein距离衡量);(ii) 从嵌入到预测的下游映射的敏感性,由Lipschitz常数刻画。二者共同构成不依赖参数数量或假设类复杂度的嵌入相关误差界。在最终嵌入层,架构敏感性消失,误差界主要由嵌入维度主导,解释了其与泛化性能之间的强经验相关性。在多种架构和数据集上的实验验证了理论,并展示了基于嵌入的诊断工具的有效性。

原文摘要 · Abstract (English)

Deep neural networks often generalize well despite heavy over-parameterization, challenging classical parameter-based analyses. We study generalization from a representation-centric perspective and analyze how the geometry of learned embeddings controls predictive performance for a fixed trained model. We show that population risk can be bounded by two factors: (i) the intrinsic dimension of the embedding distribution, which determines the convergence rate of empirical embedding distribution to the population distribution in Wasserstein distance, and (ii) the sensitivity of the downstream mapping from embeddings to predictions, characterized by Lipschitz constants. Together, these yield an embedding-dependent error bound that does not rely on parameter counts or hypothesis class complexity. At the final embedding layer, architectural sensitivity vanishes and the bound is dominated by embedding dimension, explaining its strong empirical correlation with generalization performance. Experiments across architectures and datasets validate the theory and demonstrate the utility of embedding-based diagnostics.

泛化分析嵌入维度理论机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。