arXiv:2510.21258cs.CLcs.AI2025-10NeurIPS被引 1

用分形几何度量大模型文本复杂性,揭示生成缺陷根源

Correlation Dimension of Auto-Regressive Large Language Models

  • 引入相关维数衡量语言模型的自相似结构复杂度
  • 发现预训练中存在三个明显阶段,且与幻觉倾向相关
  • 可检测多种生成退化现象,适用于各类自回归模型

大型语言模型在自然语言生成方面取得显著进展,但仍表现出重复、逻辑混乱等令人困惑的行为,即使在低困惑度下也是如此。这暴露出传统评估指标的局限性:它们侧重局部预测准确率,而忽视了长程结构复杂性。本文提出相关维数——一种分形几何意义上的自相似性度量,用于量化语言模型所感知的文本认知复杂度。该度量捕捉语言的层级重复结构,在统一框架中连接局部与全局特性。通过大量实验,我们发现相关维数(1)揭示了预训练过程中的三个不同阶段,(2)反映上下文依赖的复杂性,(3)指示模型产生幻觉的倾向,(4)能可靠检测多种生成退化的表现。该方法计算高效,对模型量化(低至4比特精度)鲁棒,广泛适用于自回归架构(如Transformer和Mamba),为理解大模型生成动态提供了新视角。

原文摘要 · Abstract (English)

Large language models (LLMs) have achieved remarkable progress in natural language generation, yet they continue to display puzzling behaviors -- such as repetition and incoherence -- even when exhibiting low perplexity. This highlights a key limitation of conventional evaluation metrics, which emphasize local prediction accuracy while overlooking long-range structural complexity. We introduce correlation dimension, a fractal-geometric measure of self-similarity, to quantify the epistemological complexity of text as perceived by a language model. This measure captures the hierarchical recurrence structure of language, bridging local and global properties in a unified framework. Through extensive experiments, we show that correlation dimension (1) reveals three distinct phases during pretraining, (2) reflects context-dependent complexity, (3) indicates a model's tendency toward hallucination, and (4) reliably detects multiple forms of degeneration in generated text. The method is computationally efficient, robust to model quantization (down to 4-bit precision), broadly applicable across autoregressive architectures (e.g., Transformer and Mamba), and provides fresh insight into the generative dynamics of LLMs.

语言模型分形几何生成质量复杂度度量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。