arXiv:2409.14981cs.LGcs.AI2024-09ICLR被引 26

研究神经模块如何专精于数据结构以实现系统性泛化。

On The Specialization of Neural Modules

  • 构建简化数据集空间,定义系统性泛化的数学标准。
  • 线性模块在任务分量求解中展现学习动态与专精度差异。
  • 理论结果可推广至复杂数据和非线性架构,验证模块化必要性。

许多机器学习模型旨在实现系统性泛化:通过组合过往经验的要素来推理新情境。这类模型采用组合式架构,学习专门处理任务中特定结构的模块,从而组合解决具有相似结构的新问题。尽管架构的组合性由设计保证,但模块是否能真正专精尚不确定。本文从实际系统性泛化基准出发,构建一个最小数据集空间,提出系统性的数学定义,并研究线性神经模块在解决任务分量时的学习动态。结果揭示了模块专精的困难、成功专精所需条件,以及模块化架构对实现系统性的必要性。最后,理论结论在更复杂数据集和非线性架构中得到验证,表明其具备泛化能力。

原文摘要 · Abstract (English)

A number of machine learning models have been proposed with the goal of achieving systematic generalization: the ability to reason about new situations by combining aspects of previous experiences. These models leverage compositional architectures which aim to learn specialized modules dedicated to structures in a task that can be composed to solve novel problems with similar structures. While the compositionality of these architectures is guaranteed by design, the modules specializing is not. Here we theoretically study the ability of network modules to specialize to useful structures in a dataset and achieve systematic generalization. To this end we introduce a minimal space of datasets motivated by practical systematic generalization benchmarks. From this space of datasets we present a mathematical definition of systematicity and study the learning dynamics of linear neural modules when solving components of the task. Our results shed light on the difficulty of module specialization, what is required for modules to successfully specialize, and the necessity of modular architectures to achieve systematicity. Finally, we confirm that the theoretical results in our tractable setting generalize to more complex datasets and non-linear architectures.

系统性泛化模块化学习神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。