arXiv:2601.19756cs.LGstat.ML2026-01被引 5

证明深度网络能高效学习分层数据结构。

Provable Learning of Random Hierarchy Models and Hierarchical Shallow-to-Deep Chaining

  • 通过逐层训练,利用标签信号与弱可识别特征实现分层学习。
  • 在温和条件下,深度卷积网络可高效学习随机分层模型。
  • 为深层网络的层次学习提供理论依据,适合研究深度学习机理者。

深度学习的成功常归因于深层网络对数据中分层结构的利用能力,能够逐层构建更复杂的特征。然而,尽管深度学习理论进展显著,多数优化结果仍集中于两到三层网络,对真正深层模型的层次学习理论理解仍然有限。这引出一个核心问题:能否证明基于梯度的方法和标准输入-标签对训练的深层网络能高效利用分层结构?本文研究了随机分层模型(Random Hierarchy Models)——一种由 arXiv:2307.02129 提出的上下文无关语法,被猜想能区分深层与浅层网络。我们证明,在温和条件下,深度卷积网络可通过梯度方法有效学习该函数类。证明的核心思想是:若中间层能接收来自标签的清晰信号,且相关特征具有弱可识别性,则逐层训练每个层即可实现分层学习。

原文摘要 · Abstract (English)

The empirical success of deep learning is often attributed to deep networks' ability to exploit hierarchical structure in data, constructing increasingly complex features across layers. Yet despite substantial progress in deep learning theory, most optimization results sill focus on networks with only two or three layers, leaving the theoretical understanding of hierarchical learning in genuinely deep models limited. This leads to a natural question: can we prove that deep networks, trained with gradient-based methods and standard input-label pairs, can efficiently exploit hierarchical structure? In this work, we consider Random Hierarchy Models -- a hierarchical context-free grammar introduced by arXiv:2307.02129 and conjectured to separate deep and shallow networks. We prove that, under mild conditions, a deep convolutional network can be efficiently trained to learn this function class. Our proof builds on a general observation: if intermediate layers can receive clean signal from the labels and the relevant features are weakly identifiable, then layerwise training each individual layer suffices to hierarchically learn the target function.

深度学习分层结构理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。