深度学习成功源于对组合稀疏结构的利用,这解释了为何超参数模型在高维任务中表现优异。
Position: A Theory of Deep Learning Must Include Compositional Sparsity
- 提出组合稀疏性是深度网络成功的根本机制
- 所有可有效图灵计算函数均具此特性,故广泛存在于现实问题中
- 适合关注理论原理与智能本质的研究者
过参数化深度神经网络(DNNs)在众多高维领域表现出色,超越了传统浅层网络受维度诅咒限制的范围。本文主张,DNNs 的成功源于其能够利用目标函数的组合稀疏结构:大多数实际相关函数可由少量基础函数组合而成,每项仅依赖输入的低维子集。我们指出,这一性质适用于所有可有效图灵计算的函数,因此极大概率存在于当前所有学习任务中。尽管已有部分理论研究揭示了组合稀疏函数下的逼近与泛化特性,但关于 DNNs 的可学习性与优化仍存关键问题。完善组合稀疏性在深度学习中的作用,是构建人工智能乃至通用智能全面理论的核心。
原文摘要 · Abstract (English)
Overparametrized Deep Neural Networks (DNNs) have demonstrated remarkable success in a wide variety of domains too high-dimensional for classical shallow networks subject to the curse of dimensionality. However, open questions about fundamental principles, that govern the learning dynamics of DNNs, remain. In this position paper we argue that it is the ability of DNNs to exploit the compositionally sparse structure of the target function driving their success. As such, DNNs can leverage the property that most practically relevant functions can be composed from a small set of constituent functions, each of which relies only on a low-dimensional subset of all inputs. We show that this property is shared by all efficiently Turing-computable functions and is therefore highly likely present in all current learning problems. While some promising theoretical insights on questions concerned with approximation and generalization exist in the setting of compositionally sparse functions, several important questions on the learnability and optimization of DNNs remain. Completing the picture of the role of compositional sparsity in deep learning is essential to a comprehensive theory of artificial, and even general, intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。