Meta-learning教会神经网络认知原语,而非简单性先验。
Meta-Learning Neural Mechanisms rather than Bayesian Priors
- 通过形式语言训练,模型学会计数等神经机制
- 仅用一种语言训练效果可比5000种语言
- 适合研究高效学习与认知建模的学者
儿童在远少于大语言模型所需数据量的情况下掌握语言。元学习被提出用于将人类学习偏见融入神经网络架构,结合符号模型的结构化泛化与神经网络的可扩展性。但元学习究竟赋予模型什么?我们研究了形式语言的元学习,发现与先前观点相反,当在以简洁性组织的数据集上训练时,元训练模型并未学习到基于简单的先验。相反,我们发现元训练会将计数等神经机制植入模型中,这些机制在下游任务中充当网络的认知原语。最令人惊讶的是,只要形式语言能激励学习有效神经机制,仅用一种语言训练的效果就相当于在5000种不同形式语言上训练。我们的发现为高效的元学习范式提供了实际启示,并为连接符号理论与神经机制提供了新的理论见解。
原文摘要 · Abstract (English)
Children acquire language despite being exposed to several orders of magnitude less data than large language models require. Meta-learning has been proposed as a way to integrate human-like learning biases into neural-network architectures, combining both the structured generalizations of symbolic models with the scalability of neural-network models. But what does meta-learning exactly imbue the model with? We investigate the meta-learning of formal languages and find that, contrary to previous claims, meta-trained models are not learning simplicity-based priors when meta-trained on datasets organised around simplicity. Rather, we find evidence that meta-training imprints neural mechanisms (such as counters) into the model, which function like cognitive primitives for the network on downstream tasks. Most surprisingly, we find that meta-training on a single formal language can provide as much improvement to a model as meta-training on 5000 different formal languages, provided that the formal language incentivizes the learning of useful neural mechanisms. Taken together, our findings provide practical implications for efficient meta-learning paradigms and new theoretical insights into linking symbolic theories and neural mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。