探究认知系统中哪些能力可自涌现,哪些必须显式计算。
Where Cognition Lives: Dissecting Emergent from Computed Function in a Minimal Complete Cognitive Architecture

- 构建最小完整认知架构,含自适应停止与价值模块。
- 停止策略看似涌现,实为仪器效应;价值计算不可替代。
- 适合研究认知机制、模型决策可解释性的学者参考。
认知架构不仅包含推理模块,还需决定思考时长与投入精力的分配。本文构建了一个最小但完整的系统:带有自适应终止的循环推理器、稳态控制场和价值模块,并检验各部分功能是通过梯度下降涌现,还是必须显式计算。结果显示,能力可涌现,停止策略看似也涌现,且其收益高于预先决定的策略(均匀策略0.467 → 难度加权0.546 → 事前价值0.698),但后续基于后验观察的收益(0.921)在审计后不成立。采用PonderNet式停止的混合隐藏状态优于固定深度基线,但当读出层均衡后,其优势消失(残差+0.000)。价值模块无法涌现:训练耦合仅捕获了显式分配器收益的零头(+0.151,路由相关性+0.79),说明在价值与内容正交的情况下,二阶决策必须显式计算。在冻结的LLM执行器上,自洽投票为可信边界(+0.0236),而样本间一致性几乎无效,其质量集中于错误答案。所有否定结论均附带机制与对照实验,协议本身亦为贡献之一。验证自身可证伪预测:承诺下的价值在悬崖代价族中获得+0.1312,是平滑族估计的七倍——非因信息提前变化,而是使可达范围扩大五倍(5.1x [3.4, 8.2])。
原文摘要 · Abstract (English)
A cognitive architecture is more than the module that reasons: it must also decide how long to think and what deserves the effort. We built a minimal but complete system - a recurrent reasoner with adaptive halting, a homeostatic control field, and a value module - and asked of each part: does this function emerge from gradient descent, or must it be computed? Competence emerges. Stopping appears to emerge too, and to be worth more than everything decidable in advance, but that appearance is instrumentation: payoff at matched mean compute climbs from 0.467 (uniform) through 0.546 (difficulty) to 0.698 (ex-ante value), and the further climb to 0.921 (posterior self-observation) does not survive audit. PonderNet-style halting returns a halting-weighted mixture of hidden states while forced-depth baselines return one, and the language head is trained on the mixture alone; equalizing the readout annihilates the apparent advantage of native execution (residual +0.000 [0.000, 0.000]). Value does not emerge: trained couplings capture zero of a payoff an explicit allocator captures completely (+0.151, routing correlation +0.79), so the second-order decisions that pay must be computed, at least where value is orthogonal to content, as here by construction. On a frozen LLM actuator the same instruments show self-consistency voting to be a measured bound (+0.0236 [+0.0150, +0.0326]) and inter-sample agreement nearly worthless as a stopping signal, its mass concentrating on wrong answers. Every null we assert carries a mechanism and a positive control, and the protocol is part of the contribution. Executing our own falsifiable prediction, value under commitment pays +0.1312 [+0.1124, +0.1502] in a cliff-cost family, some seven times the smooth-family estimate - not because the cliff shifts information ex ante, but because it multiplies the attainable range fivefold (5.1x [3.4, 8.2]).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。