发现语言模型隐空间中深层的层次结构,揭示其推理机制
H-Probes: Extracting Hierarchical Structures From Latent Representations of Language Models

- 用线性探针提取隐表示中的深度与成对距离,识别层次结构
- 在合成树遍历任务中,层次子空间低维且对高精度关键
- 真实数学推理中也存在弱层次结构,适用于理解模型思维
表征与导航层次关系是推理的基本能力。大型语言模型在多种需层次推理的任务中表现出色,但对其如何在隐空间几何上构建必要抽象结构的研究仍有限。为此,我们提出 H-probes,一组线性探针,用于从隐表示中提取层次结构,包括深度和成对距离。在合成树遍历任务中,H-probes能稳健识别完成任务所需的层次子空间;进一步的全面消融实验表明,这些层次子空间为低维,对高任务性能具有因果重要性,并在域内与域外均具泛化能力。此外,我们在真实世界层次场景(如数学推理轨迹)中也发现了类似但较弱的层次结构。结果表明,模型不仅在语法与概念层面表征层次,更在更深层抽象——包括推理过程本身——中实现层次表达。
原文摘要 · Abstract (English)
Representing and navigating hierarchy is a fundamental primitive of reasoning. Large language models have demonstrated proficiency in a wide variety of tasks requiring hierarchical reasoning, but there exists limited analysis on how the models geometrically represent the necessary latent constructions for such thinking. To this end, we develop H-probes, a collection of linear probes that extract hierarchical structure, specifically depth and pairwise distance, from latent representations. In synthetic tree traversal tasks, the H-probes robustly find the subspaces containing hierarchical structure necessary to complete the tasks; furthermore, in comprehensive ablation experiments, we show that these hierarchy-containing subspaces are low-dimensional, causally important for high task performance, and generalize within- and out-of-domain. Furthermore, we find analogous, though weaker, hierarchical structure in real-world hierarchical contexts such as mathematical reasoning traces. These results demonstrate that models represent hierarchy not only at the level of syntax and concepts, but at deeper levels of abstraction -- including the reasoning process itself.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。