arXiv:2607.04167cs.LG2026-07

语言模型用几何结构表示序数信息,局部可计算任务有清晰1维流形。

Geometry of Ordinal Representations in Language Models

论文配图:Geometry of Ordinal Representations in Language Models
图 1 · 摘自论文原文
  • 基于词元身份局部计算的序数任务形成1维流形
  • 需跨位置整合的任务产生高维或混乱表征
  • 不同架构几何变换能力差异显著,适合研究模型内部机制

近期研究表明,语言模型在弯曲的1维流形上表示字符数量,注意力头执行几何变换以实现计算。我们测试了Gemma-2-2B、Gemma-2-9B和Qwen3-4B在四种序数任务(括号深度、缩进、表格位置、数值大小)中的表现。发现当序数变量能从词元身份局部计算时,会涌现出具有位置细胞特征分块的1维流形;而需要跨位置整合或语义提取的任务则产生高维或不连贯的表征。几何计算能力依赖于架构:Qwen3-4B在缩进任务中表现出更强的扭曲,且其扭曲保持序数顺序,而数值任务中的扭曲则不具备此特性。激活修补实验确认所识别的流形子空间集中了任务相关信息,流形方向消融导致探针准确率大幅下降,远超随机方向控制组。

原文摘要 · Abstract (English)

Recent work showed that language models represent character counts on curved 1D manifolds, with attention heads performing geometric transformations to enable computation. We test whether this generalizes across four ordinal tasks (bracket depth, indentation, table position, numeric magnitude) in Gemma-2-2B, Gemma-2-9B, and Qwen3-4B. We find that 1D manifolds with place-cell feature tiling emerge for tasks where the ordinal variable is locally computable from token identity, while tasks requiring cross-position integration or semantic extraction produce higher-dimensional or incoherent representations. Geometric computation is architecture-dependent: Qwen3-4B shows substantially stronger twisting than Gemma models for indentation, and its twisters preserve ordinal order, unlike its numeric twisters. Activation patching confirms that the identified manifold subspaces concentrate task-relevant information, with manifold-direction ablation causing dramatically larger probe accuracy drops than random-direction controls.

语言模型几何表征序数推理注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。