证明了稀疏结构的神经网络能稳定学习组合特征,解释大模型为何在图像任务上有效。
Sublinearly Structured Deep Neural Networks Achieve Feature Learning Consistency for Compositional Functions

- 设计输入输出与隐藏层增长均慢于样本量的神经网络结构
- 在样本数超过参数数时仍保证特征学习一致性和预测准确性
- 解释了常见卷积网络成功的原因,适合研究深度学习理论者
过去十年中,深度神经网络(DNN)在复杂机器学习任务中取得显著成功,但其理论基础仍不完整。从统计视角看,一个自然问题是:DNN能否达到与经典模型相当的特征学习和预测一致性?尽管全面刻画尚未实现,我们为一大类模型提供了正面结果。针对输入/输出维度和隐藏神经元数量随样本量亚线性增长的子类结构,本文建立了学习分层组合目标函数时的特征学习一致性保障。重要的是,即使在传统“过参数化”情形——总参数量超过训练样本数——该一致性依然成立。实验表明,这类结构的DNN在预测性能上可媲美甚至超越宽网络。结构审计显示,广泛使用的卷积网络(如AlexNet、VGGNet、ResNet、GoogLeNet)在其图像分类任务中均属亚线性结构。进一步证明,在大样本极限下,此类DNN对分层组合函数具有普遍逼近能力。而图像本身具有内在的分层组合结构。综上所述,这些结果通过统计视角解释了大规模深度学习模型在海量图像数据上充分训练后成功的原因。
原文摘要 · Abstract (English)
Over the past decade, deep neural networks (DNNs) have achieved remarkable success on complex machine-learning tasks, yet the theoretical foundations of their performance remain incomplete. From a statistical viewpoint, a natural question is: can DNNs attain feature-learning and prediction consistency comparable to that of classical models? While a full characterization is open, we provide positive results for a broad subclass. We establish feature-learning consistency guarantees for sublinearly structured DNNs-architectures whose input/output dimensions and number of hidden neurons grow sublinearly with the sample size-when learning hierarchically compositional target functions. Importantly, this consistency still holds even in the conventional "over-parameterized" regime where the total number of parameters exceeds the number of training samples. Empirically, sublinearly structured DNNs match or surpass wide DNNs in prediction. A structural audit further indicates that widely used convolutional neural networks (CNNs), including AlexNet, VGGNet, ResNet, GoogLeNet, are sublinearly structured on their image classification benchmarks. We further prove that the sublinearly structured DNNs achieve universal approximation for hierarchically compositional functions in the large-sample limit. Moreover, images exhibit an inherent hierarchical, compositional structure. Taken together, these results explain, through a statistical lens, why many large-scale deep learning models succeed after adequate training on massive image datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。