arXiv:2505.15064cs.LGmath.DS2025-05

深度网络为何更优?研究发现:当近似能力强且状态转移几何稳定时,深度有统计优势。

Why and When Deep is Better than Shallow: Implementation-Agnostic State-Transition Model of Deep Learning

  • 用状态转移模型分析深度对泛化的影响,分离实现误差、近似误差与统计复杂度。
  • 深度依赖的方差项受球面熵积分控制,几何机制可保持其多项式增长而非指数爆炸。
  • 揭示深度优势的条件:快速近似 + 几何稳定的隐状态转移,适合理解深层结构设计。

深度为何提升泛化性能?本文在一种与实现无关的状态转移模型中研究该问题:深度为k的预测器是读出类H与由隐状态转移生成的词球B(k,F)的复合。泛化界将误差分解为实现误差、近似误差和统计复杂度,并通过覆盖熵积分对深度相关的方差项进行上界估计,同时在读出分离条件下给出下界诊断。研究识别出保持熵贡献饱和或多项式增长的几何与半群机制,对比了导致经典指数增长障碍的分离机制。结合近似率与方差上界,得出典型深度权衡模式:当近似能力快速提升而转移半群仍保持几何稳定性时,深度具有统计优势。

原文摘要 · Abstract (English)

Why and when does depth improve generalization? We study this question in an implementation-agnostic state-transition model, where a depth-$k$ predictor is a readout class $H$ composed with the word ball $B(k,F)$ generated by hidden state transitions. Generalization bounds separate implementation error, approximation error, and statistical complexity, and upper bound the depth-dependent variance term by a Dudley entropy integral over $B(k,F)$, with a conditional lower-bound diagnostic under readout separation. We identify geometric and semigroup mechanisms that keep this entropy contribution saturated or polynomial, and contrast them with separation mechanisms that recover the classical exponential-growth obstruction. Coupling these variance upper bounds with approximation rates gives typical depth trade-off patterns, clarifying that depth is statistically favorable when approximation improves rapidly while the transition semigroup remains geometrically tame.

深度学习泛化能力状态转移理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。