通过新方法揭示大模型中概念的几何结构,发现其可动态变化且支持推理。
Hypothesis-Driven Feature Manifold Analysis in LLMs via Supervised Multi-Dimensional Scaling
- 提出SMDS方法,无差别评估不同概念的几何形状。
- 发现时间推理中概念形成圆、线、聚类等不同结构。
- 适合研究模型内部表征或推理机制的学者阅读。
线性表示假设认为语言模型将概念编码为潜在空间中的方向,形成有组织的多维流形。以往研究多聚焦于单个特征的具体几何形态,难以推广。本文提出一种与模型无关的监督多维缩放(SMDS)方法,用于评估和比较不同特征流形假设。以时间推理为例,SMDS发现不同特征呈现出圆形、直线和聚类等不同的几何结构。这些结构具有稳定的语义属性,跨模型家族和规模保持一致,主动支持推理,并随上下文动态调整。结果表明,特征流形在模型中扮演功能性角色,支持基于实体的推理,即大模型通过结构化表示进行概念编码与变换。
原文摘要 · Abstract (English)
The linear representation hypothesis states that language models (LMs) encode concepts as directions in their latent space, forming organized, multidimensional manifolds. Prior work has largely focused on identifying specific geometries for individual features, limiting its ability to generalize. We introduce Supervised Multi-Dimensional Scaling (SMDS), a model-agnostic method for evaluating and comparing competing feature manifold hypotheses. We apply SMDS to temporal reasoning as a case study and find that different features instantiate distinct geometric structures, including circles, lines, and clusters. SMDS reveals several consistent characteristics of these structures: they reflect the semantic properties of the concepts they represent, remain stable across model families and sizes, actively support reasoning, and dynamically reshape in response to contextual changes. Together, our findings shed light on the functional role of feature manifolds, supporting a model of entity-based reasoning in which LMs encode and transform structured representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。