揭示大模型推理时的不确定性本质,区分真假认知误差
Descriptive versus Regulatory Uncertainty in Bounded Predictive Systems
- 提出描述性与调控性不确定性的结构区分
- 实证发现熵值恒定而准确率波动,二者无关
- 指出当前模型无法实现真正认知反馈,需物理耦合
任何在有限表征能力下建模世界系统都必须压缩信息;任何压缩都会引入先验,即系统的偏差。尚未明确的是:不确定性是否参与驱动未来行为的动态过程,还是仅描述输出分布且无实际影响。本文提出结构性区分:描述性不确定性不递归调节系统策略,而调控性不确定性直接进入优化景观并推动持续自适应重构。形式化证明当前Transformer架构在推理时受限于描述性不确定性。通过兰道尔原理将此与热力学关联:若不确定性具有调控作用,则认知误差必须消耗真实能量;在解耦系统中,幻觉与正确推导耗能相同。在三个本地部署的语言模型(3B、8B、70B参数)上进行实验,跨任务(模式检索、因果操作应用、分布外因果泛化)的分词级香农熵在所有模型中统计不变(所有成对p ≥ 0.568;单模型范围0.011-0.028纳特),而任务准确率差异显著(0%-100%)。熵与准确率正交。该解耦现象具有尺度不变性:更大模型准确率更高,但熵平坦性一致。此结构性缺陷无法通过增加参数或训练数据解决。真正的认知基础要求热力学基底状态与信息处理代价之间存在物理耦合。
原文摘要 · Abstract (English)
Any system that models the world under finite representational capacity must compress; any compression entails a prior; and the prior is the system's bias. What has not been established is whether uncertainty participates in the dynamics governing future behavior, or merely describes the output distribution without consequence. We introduce a structural distinction between descriptive uncertainty, which does not recursively modulate the system's policy, and regulatory uncertainty, which directly enters the optimization landscape and drives persistent adaptive restructuring. We prove formally that current transformer architectures are confined to descriptive uncertainty at inference. We ground this in thermodynamics via Landauer's principle: for uncertainty to be regulatory, epistemic error must cost real energy; in a decoupled system, hallucinations and correct derivations dissipate identical energy. We test this empirically across three locally-deployed language models (3B, 8B, 70B parameters). Token-level Shannon entropy is statistically invariant across tasks spanning pattern retrieval, causal operator application, and out-of-distribution causal generalization in all three models (all pairwise p >= 0.568; within-model ranges 0.011-0.028 nats), while task accuracy varies substantially across the same conditions (0%-100%). Entropy and accuracy are orthogonal. The decoupling is scale-invariant: larger models achieve higher accuracy but identical entropy flatness. This structural incapacity is not resolvable by additional parameters or training data. Genuine epistemic grounding requires physical coupling between thermodynamic substrate state and information processing cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。