首次系统分析神经网络架构复杂度演化,发现重大创新与复杂度跃升相关。
On the Architectural Complexity of Neural Networks

- 构建基于张量运算的统一理论框架,显式建模网络底层结构
- 分析40年模型演进,揭示突破性架构与复杂度增长的关联
- 生成3000+高复杂度新架构并开源,供研究者探索
我们提出一个统一的理论框架,用于对深度神经网络(DNN)进行严格分析与系统构造。该框架弥补了现有理论的空白,明确建模了张量运算的结构——这一常被抽象化的底层信息。该框架支持两项新目标:(1) 分析深度学习历史中架构复杂度的演变过程;(2) 基于新型张量运算自动生成新颖架构。我们对过去40年提出的DNN进行研究,发现突破性架构与不同类型的架构复杂度提升存在关联。此外,我们识别出多个尚未探索的高复杂度架构类别,并收集了一个包含3000多个高复杂度架构的数据集,已公开发布于:https://github.com/combinatoriallabs/ArchitecturalComplexity。
原文摘要 · Abstract (English)
We introduce a unified theoretical framework for the rigorous analysis and systematic construction of deep neural networks (DNNs). This framework addresses a gap in existing theory by explicitly modeling the structure of tensor operations -- lower level information that is often abstracted. Our framework enables two novel objectives: (1) analysis of the evolution of architectural complexity over deep learning history, and (2) automatic construction of novel architectures based on new types of tensor operations. Our study of DNNs introduced over the past 40 years reveals a connection between groundbreaking architectures and increases in different types of architectural complexity. Moreover, we identify several large classes of higher complexity architectures that have not yet been explored. We then collect a dataset of 3,000+ higher complexity architectures, which we publicly release at: https://github.com/combinatoriallabs/ArchitecturalComplexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。