arXiv:2606.30191cs.AIcs.LG2026-06被引 1

让智能体学会自我塑造行为的关键,是自因信用的缓慢积累。

From Detecting Agency to Doing Work: Self-Caused Credit Builds a Durable Behavioral Self in a Minimal Spiking Agent

论文配图:From Detecting Agency to Doing Work: Self-Caused Credit Builds a Durable Behavioral Self in a Minimal Spiking Agent
图 1 · 摘自论文原文
  • 用自因信用触发慢速参数更新,实现行为残留
  • 移除缓冲区后仍保留96%自保行为,验证持久性
  • 适合研究意识基础或持续学习的神经机制

一个能区分自我与世界的智能体,如何被这一区分持久塑造?现有研究显示预测系统可检测自身代理性(Ye, 2026),但检测不等于形成持久自我行为。本文提出:代理门控的慢速信用机制(即 Own*Agency*Salience)驱动慢参数更新,在脉冲神经网络(Nengo LIF/PES)中实现行为残留——即使在移除情景缓冲区后,自保行为仍保留96%(N=50),而重置慢解码器或关闭代理门控时则迅速崩溃。通过重构代理比较器并仅切换慢信用通道,发现仅当自因信用执行慢工作时才产生持久行为(卸载后保留率1.00 vs 0.00)。该现象在24维部分可观测控制任务中亦成立(0.74 vs 0.00),且盆地变形量等于净自因信用工作量。在八次连续任务中,乘法抑制机制有效防止遗忘(卸载后准确率达0.88,遗忘率0.13),而加法聚合、无代理对照组及无回放基线均退化至随机水平,且无需回放缓冲或任务边界保护机制。本文将这种持久残留定义为操作意义上的行为自我,主张‘自因信用完成慢工作’是构建具有自我意识代理的必要基石。不涉及意识宣称。

原文摘要 · Abstract (English)

How does an agent that can tell self from world come to be durably shaped by that distinction? Recent work shows that a predictive system can detect its own agency (Ye, 2026), but detecting agency does not explain durable, self-shaped behavior. We show that agency-gated slow credit -- a conjunctive term Own*Agency*Salience driving a slow parameter update -- produces post-unload behavioral residue: on a spiking substrate (Nengo LIF/PES), a learned self-preserving choice survives episodic buffer removal (retained fraction 0.96, N=50) and collapses when the slow decoders are reset or the agency gate is removed. Reproducing the agency comparator and toggling only the slow-credit channel, we find a clean dissociation: at matched agency gain, durable behavior develops only when self-credit performs slow work (post-unload self-preservation 1.00 vs 0.00). The same dissociation holds in 24-dimensional partially-observed control (0.74 vs 0.00), and a plastic-work analysis shows that basin deformation equals net self-credit work. Across eight sequentially-learned tasks under exogenous interference, the multiplicative veto also prevents forgetting: it retains old tasks (final post-unload accuracy 0.88, forgetting 0.13) where additive pooling collapses to chance-level recall, the no-agency ablation falls below chance, and episodic/replay baselines stay near chance after unload -- all with no replay buffer and no task-boundary-dependent protection mechanism (N=50). We formalize the durable residue as an operational behavioral self and argue that self-caused credit doing slow work is a necessary building block for agents that develop a self. No claim of consciousness is made.

神经认知自我建模持续学习脉冲网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。