用压缩率区分大模型与人类文本的统计规律差异。
The Statistical Signature of LLMs
- 通过无损压缩衡量文本结构规律性,不依赖模型内部。
- 大模型文本比人类文本更规则、更易压缩,尤其在可控场景中。
- 该方法对模型和任务均有效,适合检测生成内容真实性。
大型语言模型通过高维分布的概率采样生成文本,但这一过程如何重塑语言的结构性统计特征仍不明确。本文表明,无损压缩提供了一种简单且模型无关的统计规律度量方式,可直接从表面文本区分生成模式。我们在三个逐步复杂的语境中分析压缩行为:受控的人类-模型续写、知识基础设施的生成中介(维基百科 vs. Grokipedia),以及完全合成的社会互动环境(Moltbook vs. Reddit)。在所有场景中,压缩结果揭示了概率生成的持久结构特征。在受控和中介情境中,大模型生成的文本比人类写作具有更高的结构规律性和可压缩性,反映出输出集中于高度重复的统计模式。然而,这种特征随规模变化:在碎片化交互环境中差异减弱,表明小尺度下表面可区分性存在根本限制。该基于压缩的分离效应在不同模型、任务和领域中一致出现,仅需表面文本即可观测,无需依赖模型内部或语义评估。整体而言,研究提出一种简单而稳健的框架,量化生成系统如何重塑文本生产,为通信复杂性的演变提供了结构视角。
原文摘要 · Abstract (English)
Large language models generate text through probabilistic sampling from high-dimensional distributions, yet how this process reshapes the structural statistical organization of language remains incompletely characterized. Here we show that lossless compression provides a simple, model-agnostic measure of statistical regularity that differentiates generative regimes directly from surface text. We analyze compression behavior across three progressively more complex information ecosystems: controlled human-LLM continuations, generative mediation of a knowledge infrastructure (Wikipedia vs. Grokipedia), and fully synthetic social interaction environments (Moltbook vs. Reddit). Across settings, compression reveals a persistent structural signature of probabilistic generation. In controlled and mediated contexts, LLM-produced language exhibits higher structural regularity and compressibility than human-written text, consistent with a concentration of output within highly recurrent statistical patterns. However, this signature shows scale dependence: in fragmented interaction environments the separation attenuates, suggesting a fundamental limit to surface-level distinguishability at small scales. This compressibility-based separation emerges consistently across models, tasks, and domains and can be observed directly from surface text without relying on model internals or semantic evaluation. Overall, our findings introduce a simple and robust framework for quantifying how generative systems reshape textual production, offering a structural perspective on the evolving complexity of communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。