探索最小神经网络如何从预测世界发展出自我的因果意识。
From Prediction to Self: Developmental Conditions for Agency in Minimal Neural Systems
- 逐步构建192维GRU系统,验证四个关键条件促成自我意识形成。
- 自因预测优势(A)达94.9%,远超对照组的53.9%,证明自我表征的因果基础。
- 揭示隐性因果使用与显性自我表征可分离,适用于研究具身智能系统。
一个仅能预测世界的系统如何区分自身因果影响与外部变化?我们通过6个实验阶段、12种被证伪的替代方案及跨信号验证,在一个192维的GRU中追踪这一转变。从无行为和自我表征开始,逐项添加组件,观察系统是否能区分自因与他因变化。核心发现为‘编码鸿沟’:系统可完美补偿自身动作进行预测,却无法将‘我在行动’编码为可读状态——隐性因果使用与显性自我表征是分离能力。当满足四个条件时,该鸿沟被跨越:(1) 持久状态形成稳定吸引子,(2) 输出反馈回输入构成因果动作环,(3) 本体感觉反馈使隐性因果知识显性化,(4) 异步觉醒——先巩固感知学习再学习动作,此配置对超参数选择最鲁棒。我们提出代理增益(A = Err_world - Err_self),作为跨信号类型通用的连续度量。关键测试表明:移除外部训练信号后,因果代理的自我表征仍保持94.9%,而统计匹配对照组降至53.9%。自我表征仅在对预测有因果价值时持续存在,是因果环的内在属性,非训练产物。
原文摘要 · Abstract (English)
How does a system that merely predicts the world come to distinguish its own causal influence from everything else? We trace this transition in a minimal 192-dimensional GRU through a developmental sequence -- 6 experimental stages, 12 falsified alternatives, and cross-signal validation. Starting with no action or self-representation, we add components one at a time, tracking whether the system distinguishes self-caused from world-caused changes. The central finding is the encoding gap: a system can perfectly compensate for its own actions in prediction while failing to encode "I am acting" as a readable state -- implicit causal use and explicit self-representation are dissociated capabilities. The developmental path crosses this gap when four conditions are jointly satisfied: (1) persistent state that forms stable attractors, (2) a causal action loop linking the system's output to its input, (3) proprioceptive feedback that makes implicit causal knowledge explicit, and (4) asynchronous awakening -- consolidating perceptual learning before action learning, which yields the only configuration robust to hyperparameter choice. We propose agency gain (A = Err_world - Err_self), the predictive advantage of knowing one's own action, as a continuous metric that generalizes across signal types. A decisive test confirms the causal grounding of the encoding: after the external training signal is removed, the causal agent retains its self-representation at 94.9% while a statistically-matched control collapses to 53.9%. Self-representation persists only when causally useful for prediction -- an intrinsic property of the causal loop, not a training artifact.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。