arXiv:2510.13850cs.CLcs.AI2025-10

发现大模型推理信息密度不均,挑战了机器类人推理的假设。

Revisiting the UID Hypothesis in LLM Reasoning Traces

  • 用熵衡量推理过程的信息密度变化
  • 正确解题时信息密度波动剧烈,非均匀分布
  • 为可解释性推理模型设计提供新思路

大型语言模型(LLMs)常通过逐步推理(Chain-of-Thought, CoT)解决问题,但中间步骤往往不忠实或难以理解。受心理语言学中均匀信息密度(UID)假说启发——该假说认为人类沟通中维持稳定的信息流——我们引入基于熵的度量来分析推理轨迹中的信息流动。令人意外的是,在三个具有挑战性的数学基准测试中,我们发现大模型成功推理的整体信息流并非均匀:正确解题路径表现出剧烈的信息密度波动,与人类交流模式截然相反。这一结果挑战了关于机器推理应模仿人类的假设,并提示需重新思考可解释且自适应推理模型的设计方向。

原文摘要 · Abstract (English)

Large language models (LLMs) often solve problems using step-by-step Chain-of-Thought (CoT) reasoning, yet these intermediate steps are frequently unfaithful or hard to interpret. Inspired by the Uniform Information Density (UID) hypothesis in psycholinguistics -- which posits that humans communicate by maintaining a stable flow of information -- we introduce entropy-based metrics to analyze the information flow within reasoning traces. Surprisingly, across three challenging mathematical benchmarks, we find that successful reasoning in LLMs is globally non-uniform: correct solutions are characterized by uneven swings in information density, in stark contrast to human communication patterns. This result challenges assumptions about machine reasoning and suggests new directions for designing interpretable and adaptive reasoning models.

大模型推理信息密度可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。