arXiv:2607.25184cs.CL2026-07

发现人类语言中上下文影响随距离衰减呈1/d规律,揭示语言结构的深层统一性。

A scaling law of contextual persistence in human language

  • 用大模型探测词序影响,量化上下文持久性函数P(d)
  • P(d)在10个语料库中均近似按1/d衰减,平均指数α=1.04
  • 该规律仅存在于自然语言,对基因序列无效,适合语言学与认知研究者

人类语言在词汇层面(频率、词汇量增长)和词对层面(跨距离共现)表现出规律性结构。本文发现,词序这一决定语义的核心因素也遵循类似规律。利用大语言模型作为概率探针,测量了在词序保持原样但上下文被随机打乱时,目标困惑度的下降量;该差值即为上下文持久性函数P(d),用于分离词序的影响。在跨越六个语族、涵盖书面与口语的十个语料库中,P(d)近似按1/d衰减(P(d) ∝ d^{-α},平均α=1.04;中位数决定系数r²=0.96)。该效应在随机打乱和合成控制中消失,且在领域内模型下不适用于基因或蛋白质序列。接近1的指数表明上下文影响在对数时间尺度上大致均匀分布。结果确立了人类语言中上下文持久性的标度律。

原文摘要 · Abstract (English)

Human language exhibits lawful structure at the level of words (frequency, vocabulary growth) and word pairs (co-occurrence across distance). Here we show that the arrangement of words in sequence -- a central determinant of meaning -- obeys a comparable law. Using large language models as probabilistic probes, we measured the reduction in target perplexity conferred by prior context at distance d beyond that of the same words scrambled; this difference, the contextual persistence function P(d), isolates the influence of arrangement. Across ten corpora spanning six language families and written and spoken modalities, P(d) decayed approximately as 1/d ($P(d) \propto d^{-α}$, mean $α= 1.04$; median $r^2 = 0.96$). The effect vanished in scrambled and synthetic controls, replicated across independent probes, and did not appear in genomic or protein sequences under domain-native models. An exponent near 1 distributes contextual influence approximately uniformly across logarithmic timescales. The results establish a scaling law of contextual persistence in human language.

语言模型上下文标度律语言结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。