用神经科学方法揭示大模型内部动态演化机制
Transformer Dynamics: A neuroscientific approach to interpretability of large language models
- 将Transformer残差流视为随层演化的动力系统
- 发现单元激活跨层连续且呈加速增长趋势
- 适合对模型内部机制感兴趣的科研人员
随着人工智能模型规模与能力的爆发式增长,理解其内部机制仍是一个关键挑战。受神经科学中动力系统方法成功的启发,本文提出一种研究深度学习系统计算的新框架。聚焦于Transformer模型中的残差流(RS),将其视为随层数演化的动力系统。我们发现,尽管残差流并非特殊基,其各单元激活在层间仍表现出强连续性;激活值随层加深而加速增长并趋于密集,单个单元轨迹呈现不稳定周期轨道特征。在低维空间中,残差流沿曲线路径演化,底层具类似吸引子的动力学行为。这些发现将动力系统理论与机械可解释性相连接,为构建融合理论严谨性与大规模数据分析的‘人工智能神经科学’奠定基础,推动对现代神经网络的理解。
原文摘要 · Abstract (English)
As artificial intelligence models have exploded in scale and capability, understanding of their internal mechanisms remains a critical challenge. Inspired by the success of dynamical systems approaches in neuroscience, here we propose a novel framework for studying computations in deep learning systems. We focus on the residual stream (RS) in transformer models, conceptualizing it as a dynamical system evolving across layers. We find that activations of individual RS units exhibit strong continuity across layers, despite the RS being a non-privileged basis. Activations in the RS accelerate and grow denser over layers, while individual units trace unstable periodic orbits. In reduced-dimensional spaces, the RS follows a curved trajectory with attractor-like dynamics in the lower layers. These insights bridge dynamical systems theory and mechanistic interpretability, establishing a foundation for a "neuroscience of AI" that combines theoretical rigor with large-scale data analysis to advance our understanding of modern neural networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。