模块化神经网络在高维任务中更易泛化,且无需随维度增加样本量。
Breaking Neural Network Scaling Laws with Modularity
- 用模块化结构设计网络,可突破传统网络的泛化瓶颈。
- 在高维任务上,模块化网络所需训练数据与维度无关。
- 新学习规则提升模型在分布内外的表现,适合复杂任务建模。
模块化神经网络在视觉问答到机器人控制等任务中表现优于非模块化网络,其优势源于对现实问题组合结构的更好建模。然而,关于模块化如何提升泛化能力,以及训练时如何利用任务模块性仍缺乏理论解释。基于近期神经网络泛化理论,我们研究了任务输入内在维度与所需训练数据量的关系。理论上证明:对于模块化任务,非模块化网络需要随任务维度呈指数增长的样本数,而模块化网络的样本复杂度与维度无关,可在高维下实现泛化。随后,我们提出一种新型学习规则以利用此优势,并在高维模块化任务上实证验证了该规则在分布内和分布外均显著提升泛化性能。
原文摘要 · Abstract (English)
Modular neural networks outperform nonmodular neural networks on tasks ranging from visual question answering to robotics. These performance improvements are thought to be due to modular networks' superior ability to model the compositional and combinatorial structure of real-world problems. However, a theoretical explanation of how modularity improves generalizability, and how to leverage task modularity while training networks remains elusive. Using recent theoretical progress in explaining neural network generalization, we investigate how the amount of training data required to generalize on a task varies with the intrinsic dimensionality of a task's input. We show theoretically that when applied to modularly structured tasks, while nonmodular networks require an exponential number of samples with task dimensionality, modular networks' sample complexity is independent of task dimensionality: modular networks can generalize in high dimensions. We then develop a novel learning rule for modular networks to exploit this advantage and empirically show the improved generalization of the rule, both in- and out-of-distribution, on high-dimensional, modular tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。