用神经科学方法分析大模型生成时的动态规律,发现不同任务下有明显差异。
Dynamical Systems Analysis Reveals Functional Regimes in Large Language Models
- 借鉴神经科学中的动态整合与亚稳态概念,构建用于评估语言模型内部动态的指标。
- 在五种不同条件下测试,结构化推理时动态指标显著高于重复、噪声和扰动情形。
- 该方法对层、通道和随机种子不敏感,适合研究模型内部计算机制差异。
大型语言模型通过高维内部动态进行文本生成,但其时间结构仍不清楚。现有解释性方法多关注静态表征或因果干预,忽视了时间动态。受神经科学启发,我们引入时间整合与亚稳态作为核心指标,基于自回归生成过程中的激活时序数据计算复合动态度量。在 GPT-2-medium 上评估五种条件:结构化推理、强制重复、高温噪声采样、注意力头剪枝和权重噪声注入。结果显示,结构化推理下的动态指标显著高于重复、噪声及扰动场景,经单因素方差分析验证,关键对比中效应量较大。结果对层选择、通道子采样和随机种子均具鲁棒性。研究证明,基于神经科学的动态度量可有效刻画大模型在不同功能模式下的计算组织差异。该度量反映形式化的动态特性,不暗示主观意识。
原文摘要 · Abstract (English)
Large language models perform text generation through high-dimensional internal dynamics, yet the temporal organisation of these dynamics remains poorly understood. Most interpretability approaches emphasise static representations or causal interventions, leaving temporal structure largely unexplored. Drawing on neuroscience, where temporal integration and metastability are core markers of neural organisation, we adapt these concepts to transformer models and discuss a composite dynamical metric, computed from activation time-series during autoregressive generation. We evaluate this metric in GPT-2-medium across five conditions: structured reasoning, forced repetition, high-temperature noisy sampling, attention-head pruning, and weight-noise injection. Structured reasoning consistently exhibits elevated metric relative to repetitive, noisy, and perturbed regimes, with statistically significant differences confirmed by one-way ANOVA and large effect sizes in key comparisons. These results are robust to layer selection, channel subsampling, and random seeds. Our findings demonstrate that neuroscience-inspired dynamical metrics can reliably characterise differences in computational organisation across functional regimes in large language models. We stress that the proposed metric captures formal dynamical properties and does not imply subjective experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。