用简单模型揭示深度学习中反直觉现象的内在机制。
Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond
- 构建逐层近似串联的简化模型,捕捉网络训练本质。
- 实证发现该模型可预测双下降、突现学习等异常性能。
- 适合研究者理解训练机制,尤其关注优化与架构设计者。
深度学习时常表现出意料之外的行为。为深入理解这些反直觉现象,本文提出一种由一系列一阶近似串联而成的简洁而精确的神经网络模型,可作为实用分析工具。通过三个案例研究,展示了该模型在解释文献中的多种显著现象——包括双下降、突现学习(grokking)、线性模式连通性,以及在表格数据上应用深度学习的挑战——并证明其能构造出可预测和解释神经网络意外表现的度量指标。此外,该模型还提供了一种教学性形式化框架,即使在复杂现代训练场景中也能分离训练过程各组件,帮助分析架构与优化策略的影响,并揭示了神经网络学习与梯度提升之间令人惊讶的相似性。
原文摘要 · Abstract (English)
Deep learning sometimes appears to work in unexpected ways. In pursuit of a deeper understanding of its surprising behaviors, we investigate the utility of a simple yet accurate model of a trained neural network consisting of a sequence of first-order approximations telescoping out into a single empirically operational tool for practical analysis. Across three case studies, we illustrate how it can be applied to derive new empirical insights on a diverse range of prominent phenomena in the literature -- including double descent, grokking, linear mode connectivity, and the challenges of applying deep learning on tabular data -- highlighting that this model allows us to construct and extract metrics that help predict and understand the a priori unexpected performance of neural networks. We also demonstrate that this model presents a pedagogical formalism allowing us to isolate components of the training process even in complex contemporary settings, providing a lens to reason about the effects of design choices such as architecture & optimization strategy, and reveals surprising parallels between neural network learning and gradient boosting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。