LLM推理时信息流不均匀反而更优,与人类沟通习惯相反。
Revisiting the Uniform Information Density Hypothesis in LLM Reasoning
- 用熵度量信息密度,分层次评估推理过程
- 高质量推理具局部平滑、全局非均匀特征
- 揭示模型推理与人类沟通的根本差异
统一信息密度(UID)假说认为,有效沟通需保持信息流稳定。本文在大语言模型(LLM)推理背景下重新审视该原则,探究步骤级均匀性是否反映推理质量。为此,我们提出一种新框架,基于熵的逐步密度度量,在局部和全局层面量化信息流均匀性。在七个推理基准上的实验显示反直觉现象:高质量推理呈现局部均匀性与全局非均匀性共存的特征。结果表明,此类均匀性优于其他内部信号,可更准确预测推理质量。这一与人类沟通模式的差异并非模型缺陷,而是源于人类沟通与模型推理目标的本质不同。
原文摘要 · Abstract (English)
The Uniform Information Density (UID) hypothesis proposes that effective communication is achieved by maintaining a stable flow of information. In this work, we revisit this principle in the context of Large Language Model (LLM) reasoning, asking whether step-level uniformity reflects reasoning quality. To this end, we introduce a novel framework to quantify uniformity of information flow at both local and global levels, using an entropy-based stepwise density metric. Across experiments on seven reasoning benchmarks, we see a counter-intuitive pattern: while high-quality reasoning exhibit smooth step-by-step transitions local uniformity and structured, non-uniform information flow at the trajectory level global non-uniformity. The results demonstrate that these uniformities outperform alternative internal signals as predictors of reasoning quality, and such divergence with human communication is not a model deficiency, but a byproduct of distinct objectives between human communication and LLM reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。