arXiv:2409.09626cs.LGcs.AI2024-09被引 7

发现组合映射最简,解释模型为何能良好泛化。

Understanding Simplicity Bias towards Compositional Mappings via Learning Dynamics

  • 用编码长度衡量,组合映射是最简单的双射
  • 神经网络训练天然倾向学习这类简单映射
  • 适合研究模型泛化机制的学者阅读

获得组合映射对于模型实现良好的组合泛化至关重要。为更好理解何时以及如何促使模型学习此类映射,本文从多个角度研究其唯一性。首先,我们证明组合映射是通过编码长度(即柯尔莫戈洛夫复杂度的上界)衡量下的最简单双射,这一性质解释了为何具备此类映射的模型能够实现良好泛化。进一步地,我们表明这种简化偏置通常是梯度下降训练下神经网络的内在特性,这在一定程度上解释了为何某些模型在适当训练后能自发实现良好泛化。

原文摘要 · Abstract (English)

Obtaining compositional mappings is important for the model to generalize well compositionally. To better understand when and how to encourage the model to learn such mappings, we study their uniqueness through different perspectives. Specifically, we first show that the compositional mappings are the simplest bijections through the lens of coding length (i.e., an upper bound of their Kolmogorov complexity). This property explains why models having such mappings can generalize well. We further show that the simplicity bias is usually an intrinsic property of neural network training via gradient descent. That partially explains why some models spontaneously generalize well when they are trained appropriately.

组合泛化模型机制学习动态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。