arXiv:2410.01545cs.LGphysics.data-an2024-10ICLR被引 8

揭示大模型推理轨迹的低维统计规律。

Lines of Thought in Large Language Models

  • 通过分析模型内部路径,发现思维轨迹在低维非欧空间聚集。
  • 轨迹可由少量参数的随机方程精确近似。
  • 为理解大模型决策机制提供新视角,适合研究者参考。

大语言模型通过连续变压器层将文本向量在嵌入空间中传递,实现下一个词的预测。这一过程产生的高维轨迹对应不同的上下文化步骤,完全决定输出概率分布。我们旨在刻画这些‘思维轨迹’的统计特性。观察发现,独立轨迹在低维非欧流形上聚集,其路径可用少量参数的随机方程良好逼近。令人惊讶的是,如此复杂的大模型行为可被简化为更简洁的形式,这对理解模型运作具有深远意义。

原文摘要 · Abstract (English)

Large Language Models achieve next-token prediction by transporting a vectorized piece of text (prompt) across an accompanying embedding space under the action of successive transformer layers. The resulting high-dimensional trajectories realize different contextualization, or 'thinking', steps, and fully determine the output probability distribution. We aim to characterize the statistical properties of ensembles of these 'lines of thought.' We observe that independent trajectories cluster along a low-dimensional, non-Euclidean manifold, and that their path can be well approximated by a stochastic equation with few parameters extracted from data. We find it remarkable that the vast complexity of such large models can be reduced to a much simpler form, and we reflect on implications.

大模型轨迹分析统计规律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。