arXiv:2607.10578cs.AI2026-07

用拉盖尔几何解析大模型概念结构,揭示其内在层次与推理轨迹。

Laguerre Geometry for Interpreting Large Language Models

论文配图:Laguerre Geometry for Interpreting Large Language Models
图 1 · 摘自论文原文
  • 提出拉盖尔几何框架,将概念定义为区域而非点或方向。
  • 发现权重可自然揭示概念的包含与层级关系,精度达100%。
  • 开发无需训练的Geometric Lens,实时读取隐藏向量编码的概念。

现有假设将大语言模型中的概念表示为单点、线性方向或高斯簇,但其形成机制尚不明确。本文表明,概念几何可通过拉盖尔几何精确刻画:概念被定义为一个区域——拉盖尔-沃罗诺伊胞或胞的并集,从而实现概念的严格定义、测量与分离。基于此,我们发现更细粒度的概念结构(如包含与层次)由拉盖尔权重自然呈现。进一步地,我们将该几何嵌入Transformer中:将每层分解为分段线性算子,揭示了标记隐藏轨迹受两个耦合机制支配——静态的分段线性流树,以及跨标记注意力触发时动态的轨迹跳跃。该分解催生了Geometric Lens,一种无需训练、无超参数的方法,可精确读取任一层隐藏向量所编码的概念。我们还构建了拉盖尔自编码器(Laguerre Autoencoder),在二维视图中同时可视化决策几何与模型完整推理轨迹。最后,我们从解释性迈向可操作性,证明Geometric Lens在上下文干扰提示下仍能准确恢复事实标记。代码已开源。

原文摘要 · Abstract (English)

Existing hypotheses represent a concept in an LLM as a single point, a linear direction, or a Gaussian cluster, yet it remains unclear how and why such structures emerge. Here, we show that concept geometry can be precisely characterized via Laguerre Geometry, in which a concept is defined as a region--a Laguerre-Voronoi cell or a union of cells--allowing us to strictly define, measure, and separate concepts. Building on this formulation, we show that finer-grained concept structures, such as inclusion and hierarchy, are naturally revealed by the Laguerre weights. We then push this geometry inside the transformer. Decomposing each layer into piecewise-linear operators, we show that a token's hidden trajectory is governed by two coupled mechanisms: a static tree of self-contained piecewise-linear flow, and a dynamic transport that hops the trajectory across trees when cross-token attention fires. This decomposition yields Geometric Lens, a training-free, hyperparameter-free method for reading out the exact concept a hidden vector encodes at any layer. We also develop Laguerre Autoencoder, a 2D visualizer that renders both the decision geometry and a model's full reasoning trajectory in one view. Finally, we move beyond explanatory geometry toward actionable interpretability, showing that Geometric Lens recovers the correct factual token when a model is prompted with in-context interference. The code is available on GitHub.

大模型解释几何建模注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。