arXiv:2508.00331cs.LG2025-08被引 2

用物理中的敏感度分析,看语言模型如何长出神经结构。

Embryology of a Language Model

  • 用UMAP可视化模型训练中敏感度矩阵,追踪结构演化。
  • 发现已知的诱导回路和未知的计数空格结构。
  • 适合研究大模型内部工作机制的科研人员。

理解语言模型如何发展其内部计算结构,是深度学习科学的核心问题。尽管来自统计物理的敏感度分析提供了有前景的分析工具,但其在可视化网络组织方面的潜力尚未被充分挖掘。本文提出一种胚胎学方法,将UMAP应用于敏感度矩阵,以可视化模型在训练过程中的结构发育。可视化结果揭示了清晰的“身体蓝图”,展现了已知特征(如诱导回路)的形成,并发现了此前未知的结构——一个专门用于计数空格标记的“间距鳍”。本工作表明,敏感度分析可超越验证阶段,用于发现新机制,为研究复杂神经网络的发育规律提供强大而全面的视角。

原文摘要 · Abstract (English)

Understanding how language models develop their internal computational structure is a central problem in the science of deep learning. While susceptibilities, drawn from statistical physics, offer a promising analytical tool, their full potential for visualizing network organization remains untapped. In this work, we introduce an embryological approach, applying UMAP to the susceptibility matrix to visualize the model's structural development over training. Our visualizations reveal the emergence of a clear ``body plan,'' charting the formation of known features like the induction circuit and discovering previously unknown structures, such as a ``spacing fin'' dedicated to counting space tokens. This work demonstrates that susceptibility analysis can move beyond validation to uncover novel mechanisms, providing a powerful, holistic lens for studying the developmental principles of complex neural networks.

语言模型神经网络可视化发育机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。