探究深度神经网络的结构如何影响性能,提出可优化的结构建模方法。
Structure of Artificial Neural Networks -- Empirical Investigations
- 基于图结构定义神经网络架构,统一搜索与优化框架。
- 实证分析结构对准确率、鲁棒性及能耗的影响,发现结构显著决定性能。
- 提出高效预测模型与生成采样策略,加速架构搜索过程。
十年间,深度学习取代了人工智能众多问题的主流解决方案。'深度'指代在无直接观测的流形上进行操作的深层架构。尽管这些架构预设了某种结构,但其具体形态尚不明确。本文提出神经网络结构的正式定义,使神经架构搜索问题与求解方法能在统一框架下建模。实践与理论之间的鸿沟由此引发:结构是否关键,还是可任意选择?本研究聚焦深度神经网络的结构,基于经验原则探索自动构建方法,以揭示所谓“黑箱模型”的内在机制。主要贡献包括:提出图诱导神经网络的形式化建模,用于定义架构优化问题;分析不同目标(如正确性、鲁棒性、能耗)下的结构特性;比较多种自动化架构优化方法的适用性。基于这些洞见,本文在两个方面推进了现有方法:一是提出替代计算成本高昂评估方案的新预测模型;二是分析并讨论了用于架构搜索中智能采样的新生成模型。
原文摘要 · Abstract (English)
Within one decade, Deep Learning overtook the dominating solution methods of countless problems of artificial intelligence. ``Deep'' refers to the deep architectures with operations in manifolds of which there are no immediate observations. For these deep architectures some kind of structure is pre-defined -- but what is this structure? With a formal definition for structures of neural networks, neural architecture search problems and solution methods can be formulated under a common framework. Both practical and theoretical questions arise from closing the gap between applied neural architecture search and learning theory. Does structure make a difference or can it be chosen arbitrarily? This work is concerned with deep structures of artificial neural networks and examines automatic construction methods under empirical principles to shed light on to the so called ``black-box models''. Our contributions include a formulation of graph-induced neural networks that is used to pose optimisation problems for neural architecture. We analyse structural properties for different neural network objectives such as correctness, robustness or energy consumption and discuss how structure affects them. Selected automation methods for neural architecture optimisation problems are discussed and empirically analysed. With the insights gained from formalising graph-induced neural networks, analysing structural properties and comparing the applicability of neural architecture search methods qualitatively and quantitatively we advance these methods in two ways. First, new predictive models are presented for replacing computationally expensive evaluation schemes, and second, new generative models for informed sampling during neural architecture search are analysed and discussed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。