arXiv:2502.20237cs.LGcs.AI2025-02被引 6

初始权重比网络结构更能决定神经网络的学习能力。

Teasing Apart Architecture and Initial Weights as Sources of Inductive Bias in Neural Networks

  • 用元学习优化初始权重,减少不同架构间的性能差异。
  • 优化后,不同网络结构在三类任务上表现趋同,差距大幅缩小。
  • 适合研究学习机制、模型偏见或元学习的学者阅读。

人工神经网络可从数据中习得人类知识的诸多方面,使其成为人类学习的潜在模型。但其学习能力受归纳偏置影响——即除数据外影响模型解的因素——而目前对神经网络的归纳偏置理解仍不充分,限制了我们基于系统表现推断人类学习的能力。认知科学与机器学习研究者通常关注网络结构作为归纳偏置来源。本文通过元学习方法,探索另一来源——初始权重的影响。我们在三种需不同偏置和泛化形式的任务上,对四种常用架构(MLPs、CNNs、LSTMs、Transformers)进行了430个模型的元训练。结果发现,元学习可显著降低甚至完全消除不同架构与数据表示间的性能差异,表明这些因素可能并非如通常认为的那样是主要的归纳偏置来源。当差异仍存在时,元学习前表现优异的架构与数据表示往往更易元训练。此外,所有架构在远离其元训练经验的问题上均表现不佳,凸显强化归纳偏置对鲁棒泛化的重要性。

原文摘要 · Abstract (English)

Artificial neural networks can acquire many aspects of human knowledge from data, making them promising as models of human learning. But what those networks can learn depends upon their inductive biases -- the factors other than the data that influence the solutions they discover -- and the inductive biases of neural networks remain poorly understood, limiting our ability to draw conclusions about human learning from the performance of these systems. Cognitive scientists and machine learning researchers often focus on the architecture of a neural network as a source of inductive bias. In this paper we explore the impact of another source of inductive bias -- the initial weights of the network -- using meta-learning as a tool for finding initial weights that are adapted for specific problems. We evaluate four widely-used architectures -- MLPs, CNNs, LSTMs, and Transformers -- by meta-training 430 different models across three tasks requiring different biases and forms of generalization. We find that meta-learning can substantially reduce or entirely eliminate performance differences across architectures and data representations, suggesting that these factors may be less important as sources of inductive bias than is typically assumed. When differences are present, architectures and data representations that perform well without meta-learning tend to meta-train more effectively. Moreover, all architectures generalize poorly on problems that are far from their meta-training experience, underscoring the need for stronger inductive biases for robust generalization.

神经网络归纳偏置元学习模型架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。