arXiv:2608.14361cs.CL2026-08

发现语言模型表征复杂度受词汇多样性影响,存在高低两种调控机制。

Local and Global Regimes of Geometric Complexity in Language Model Representations

  • 通过分析数据集末尾词的唯一性,揭示内在维度随词汇多样性变化的双态规律。
  • 低多样性时独特词少则内在维度高,高多样性时相反,转折点可精确计算。
  • 为理解大模型内部表征结构提供新视角,适合研究模型表征与数据构造关系者。

内在维度(ID)被广泛用于探测语言模型的表征复杂度,但其差异究竟是语言本身特性还是数据构建方式造成的尚不明确。本文聚焦于词汇多样性(即数据集中唯一末尾词的数量)如何影响数据集的ID估计。我们发现存在一种尺度依赖的双态转变:在低词汇多样性下,唯一末尾词较少的条件产生更高的ID;而在高词汇多样性下,这一顺序反转,唯一词较多的条件反而导致更高ID。我们推导出一个无需参数的精确公式,准确预测了该转变点,在所有测试尺度上均与实际观测一致。一方面,结果表明不能将表征的内在维度简单视为复杂度的指标;另一方面,发现的两种ID行为模式揭示了语言数据在大模型中的普遍组织原则,为理解其内部流形结构提供了新见解。

原文摘要 · Abstract (English)

Intrinsic dimensionality (ID) is widely used to probe the representational complexity of language models, but it remains unclear whether ID differences reflect properties of language itself or artefacts of how the underlying dataset was constructed. In this paper, we focus specifically on how lexical diversity, the number of unique last-token items present in a dataset, affects ID estimates of that dataset. We find a scale-dependent transition between two regimes: at low lexical diversity, conditions with fewer unique final words produce higher ID, while at high lexical diversity, this ordering reverses, and conditions with more unique words produce higher ID. We derive an exact, parameter-free formula for the point at which this reversal occurs, which matches the observed transition point at every scale tested. On the one hand, our results highlight how care must be taken when interpreting the intrinsic dimensionality of a set of representations as a straightforward cue of their complexity. On the other hand, our discovery of the two ID regimes reveals a general principle of organisation of linguistic data in LLMs that sheds new light on their inner manifold structures.

语言模型表征复杂度内在维度数据构造

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。