发现大模型残差流中存在稳定的激活区域,与语义划分相关。
Characterizing stable regions in the residual stream of LLMs
- 通过分析残差流,发现模型输出对小激活变化不敏感的稳定区域。
- 这些区域随训练推进和模型增大而更清晰,边界处敏感度高。
- 区域内部语义相近,适合研究模型可解释性与训练动态。
我们识别出Transformer模型残差流中的稳定区域:在区域内,模型输出对微小激活变化不敏感,但在区域边界处表现出高度敏感性。这些区域在训练过程中逐渐形成,并随着训练进程或模型规模增加而变得更加明确。其范围远大于此前研究中的多面体结构。分析表明,这些稳定区域与语义区分一致,相似提示在区域内聚集,同一区域内的激活产生相似的下一个词预测。该工作为理解神经网络复杂性提供了新方向,有助于揭示训练动态并推动可解释性研究。
原文摘要 · Abstract (English)
We identify stable regions in the residual stream of Transformers, where the model's output remains insensitive to small activation changes, but exhibits high sensitivity at region boundaries. These regions emerge during training and become more defined as training progresses or model size increases. The regions appear to be much larger than previously studied polytopes. Our analysis suggests that these stable regions align with semantic distinctions, where similar prompts cluster within regions, and activations from the same region lead to similar next token predictions. This work provides a promising research direction for understanding the complexity of neural networks, shedding light on training dynamics, and advancing interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。