arXiv:2604.05536cs.CLcs.AI2026-04

语言嵌入中发现类湍流5/3幂律,揭示语义的跨尺度自相似组织。

Turbulence-like 5/3 spectral scaling in contextual representations of language as a complex system

  • 将文本视为嵌入空间中的轨迹,用嵌入步信号量化词元序列的尺度波动。
  • 多语言、多语料下功率谱呈5/3幂律,覆盖广泛频率范围。
  • 该现象反映上下文依赖的多尺度结构,适合研究语言复杂性与模型表征。

自然语言是一个表现出稳健统计规律的复杂系统。本文将文本表示为基于Transformer的语言模型生成的高维嵌入空间中的轨迹,并利用嵌入步信号量化词元序列上的尺度依赖波动。在多种语言和语料上,得到的功率谱呈现出稳定的幂律,指数接近5/3,且在较宽频率范围内保持。这一特征在人类写作与AI生成文本的上下文嵌入中均一致出现,但在静态词向量中缺失,并在词元顺序随机化后被破坏。结果表明,该幂律反映了多尺度、上下文依赖的组织结构,而非仅由词汇统计决定。类比湍流中的科尔莫戈罗夫谱,研究提示语义信息以无标度、自相似的方式在语言尺度间整合,为研究语言表征中的复杂结构提供了定量、模型无关的基准。

原文摘要 · Abstract (English)

Natural language is a complex system that exhibits robust statistical regularities. Here, we represent text as a trajectory in a high-dimensional embedding space generated by transformer-based language models, and quantify scale-dependent fluctuations along the token sequence using an embedding-step signal. Across multiple languages and corpora, the resulting power spectrum exhibits a robust power law with an exponent close to $5/3$ over an extended frequency range. This scaling is observed consistently in contextual embeddings from both human-written and AI-generated text, but is absent in static word embeddings and is disrupted by randomization of token order. These results show that the observed scaling reflects multiscale, context-dependent organization rather than lexical statistics alone. By analogy with the Kolmogorov spectrum in turbulence, our findings suggest that semantic information is integrated in a scale-free, self-similar manner across linguistic scales, and provide a quantitative, model-agnostic benchmark for studying complex structure in language representations.

语言模型幂律复杂系统嵌入分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。