arXiv:2609.07053cs.CL2026-09

首次系统分析大模型隐藏状态的树状结构,发现中层最像树形网络。

A Hyperbolicity Atlas of Large Language Model Hidden States

论文配图:A Hyperbolicity Atlas of Large Language Model Hidden States
图 1 · 摘自论文原文
  • 用格罗莫夫双曲性衡量大模型各层隐藏状态的距离结构
  • 中层隐藏状态双曲性最高,末层更接近树状结构
  • 适合研究模型内部表征与结构特性的研究人员

大语言模型的隐藏状态是普通向量,但这些向量之间的距离仍可能呈现层次结构。据我们所知,本文是首个系统研究当代大模型中提示-标记隐藏状态是否表现出格罗莫夫双曲性(Gromov Hyperbolicity, GH)的研究。基于来自十款开源模型在MATH500、HumanEval、WinoGrande和TruthfulQA四个数据集上的818,904个样本-层测量数据,构建了覆盖参数规模、层数深度、模型家族和输入领域四个维度的GH图谱。最显著模式出现在深度维度:中层通常形成相对高双曲性的平台,而末层则明显更具树状特征。参数规模影响较弱且非单调,7/8B参数量级的模型家族间差异显著,不同输入领域也与模型专长产生交互作用。这些发现使双曲性成为实用诊断工具:可揭示层级距离结构出现的位置,展示模型专长如何改变其结构,并指出值得深入比较的模型-层-领域组合。

原文摘要 · Abstract (English)

LLM hidden states are ordinary vectors, but the distances among those vectors may still show hierarchical structure. To our knowledge, this paper is the first systematic study of whether prompt-token hidden states in contemporary LLMs exhibit Gromov Hyperbolicity (GH), a distance-based measure of tree-likeness. Using 818,904 sample-layer measurements from ten open-weight models across MATH500, HumanEval, WinoGrande, and TruthfulQA, we build a GH map over four axes: parameter scale, layer depth, model family, and input domain. The clearest pattern is depth, not scale: middle layers usually form a high-relative-hyperbolicity plateau, while final layers often become substantially more tree-like. Scale effects are weak and non-monotonic, matched 7/8B model families differ strongly, and domains interact with model specialization. These findings make GH useful as a practical diagnostic: it shows where hierarchical distance structure appears, how specialization changes it, and which model-layer-domain comparisons deserve closer analysis.

大模型表征双曲几何模型结构分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。