arXiv:2501.12391cs.LGcs.AI2025-01被引 8

揭示神经网络学技能的顺序规律,提出三类简化模型解释学习机制。

Physics of Skill Learning

  • 用几何、资源和多米诺三类模型抽象学习过程
  • 发现技能按序学习,类似多米诺骨牌连锁反应
  • 模型启发实际训练加速算法,适合研究学习动力学者

我们旨在理解神经网络训练中技能学习的物理机制。观察到技能呈序列式学习,且前一技能完成后会立即触发下一技能的学习,类似多米诺骨牌依次倒下。借鉴物理学家的抽象与简化方法,提出三种复杂度递增的模型——几何模型、资源模型与多米诺模型。几何模型可复现多米诺效应,其资源视角启发了资源模型,后者可进一步简化为多米诺模型。三类模型分别揭示了神经缩放定律、组合任务学习动态及模块化优势。几何模型可导出Chinchilla缩放律,资源模型解释复合任务学习,多米诺模型展现模块化价值。这些模型不仅概念上有趣,且推动实际算法改进,如基于模型设计的简单修改可显著加速深度学习训练。

原文摘要 · Abstract (English)

We aim to understand physics of skill learning, i.e., how skills are learned in neural networks during training. We start by observing the Domino effect, i.e., skills are learned sequentially, and notably, some skills kick off learning right after others complete learning, similar to the sequential fall of domino cards. To understand the Domino effect and relevant behaviors of skill learning, we take physicists' approach of abstraction and simplification. We propose three models with varying complexities -- the Geometry model, the Resource model, and the Domino model, trading between reality and simplicity. The Domino effect can be reproduced in the Geometry model, whose resource interpretation inspires the Resource model, which can be further simplified to the Domino model. These models present different levels of abstraction and simplification; each is useful to study some aspects of skill learning. The Geometry model provides interesting insights into neural scaling laws and optimizers; the Resource model sheds light on the learning dynamics of compositional tasks; the Domino model reveals the benefits of modularity. These models are not only conceptually interesting -- e.g., we show how Chinchilla scaling laws can emerge from the Geometry model, but also are useful in practice by inspiring algorithmic development -- e.g., we show how simple algorithmic changes, motivated by these toy models, can speed up the training of deep learning models.

技能学习神经网络模型简化学习动态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。