arXiv:2509.19554cs.LGcs.AI2025-09被引 1

用力学类比分析模型训练中样本间的相互影响。

Learning Dynamics of Deep Learning -- Force Analysis of Deep Neural Networks

  • 将训练样本间的影响分解为相似度与更新力两个维度。
  • 解释了某些样本学习路径非平凡、微调方法有效性的原因。
  • 适合研究模型训练机制或优化策略的学者参考。

本论文探讨深度学习模型在训练过程中的学习动态,借鉴力学中的受力分析思想。具体而言,通过聚焦模型训练过程,分析单个训练样本对另一个样本的影响,如同分析力如何推动物体运动。我们将这种影响分解为两部分:两个样本之间的相似性,以及更新力度的强弱。该框架帮助我们理解多种实际系统中模型的行为,例如解释为何某些样本具有非平凡的学习路径,为何某些大语言模型微调方法有效(或无效),以及为何更简单、结构化的模式更容易被学习。我们将该方法应用于多个学习任务,发现了改进模型训练的新策略。尽管该方法仍在发展中,但为系统性地解读模型行为提供了新视角。

原文摘要 · Abstract (English)

This thesis explores how deep learning models learn over time, using ideas inspired by force analysis. Specifically, we zoom in on the model's training procedure to see how one training example affects another during learning, like analyzing how forces move objects. We break this influence into two parts: how similar the two examples are, and how strong the updating force is. This framework helps us understand a wide range of the model's behaviors in different real systems. For example, it explains why certain examples have non-trivial learning paths, why (and why not) some LLM finetuning methods work, and why simpler, more structured patterns tend to be learned more easily. We apply this approach to various learning tasks and uncover new strategies for improving model training. While the method is still developing, it offers a new way to interpret models' behaviors systematically.

深度学习训练动态模型机制力分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。