arXiv:2411.14288cs.LGmath.ST2024-11中稿 · ed被引 1

解析卷积网络设计对泛化能力的影响,揭示权值共享与局部性的作用机制。

On the Sample Complexity of One Hidden Layer Networks with Equivariance, Locality and Weight Sharing

  • 基于统计学习理论,分析单层网络中权值共享、等变性和局部性对样本复杂度的影响。
  • 给出滤波器范数决定的维数无关上下界,且对多种激活函数适用。
  • 发现非等变权值共享也能达到类似等变效果,局部性与表达力存在权衡。

权值共享、等变性与局部滤波器是卷积神经网络提升样本效率的关键设计,但它们各自对泛化误差的贡献尚不明确。本文从统计学习理论出发,刻画这些设计选择对样本复杂度的相对影响。研究给出了单隐层网络的一类下界与上界,对一大类激活函数而言,这些界仅依赖于滤波器范数且与维度无关。还提供了最大池化及多层网络的边界,其依赖关系较弱。理论分析表明:在特定权值共享机制下,非等变共享可获得与等变共享相当的泛化界;局部性具有泛化优势,但不确定性原理暗示其与表达力之间存在权衡。通过大量实验验证了这些规律的一致性。

原文摘要 · Abstract (English)

Weight sharing, equivariance, and local filters, as in convolutional neural networks, are believed to contribute to the sample efficiency of neural networks. However, it is not clear how each one of these design choices contributes to the generalization error. Through the lens of statistical learning theory, we aim to provide insight into this question by characterizing the relative impact of each choice on the sample complexity. We obtain lower and upper sample complexity bounds for a class of single hidden layer networks. For a large class of activation functions, the bounds depend merely on the norm of filters and are dimension-independent. We also provide bounds for max-pooling and an extension to multi-layer networks, both with mild dimension dependence. We provide a few takeaways from the theoretical results. It can be shown that depending on the weight-sharing mechanism, the non-equivariant weight-sharing can yield a similar generalization bound as the equivariant one. We show that locality has generalization benefits, however the uncertainty principle implies a trade-off between locality and expressivity. We conduct extensive experiments and highlight some consistent trends for these models.

深度学习泛化理论卷积网络样本复杂度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。