arXiv:2608.11506cs.NEcs.LG2026-08

智能体在信息不全时,靠内部状态预测未来并调控行为。

Predictive Allostatic Organization in Recurrent and Spiking Agents Under Partial Observability

  • 用递归和脉冲神经网络模拟智能体,学习在部分可观测环境中生存。
  • 内部动态可提前预测成功获取资源与安全行为,最高准确率达0.802。
  • 适合研究具身智能、预测性控制及能量敏感决策的学者。

在资源获取、避险、接触依赖性摄食及内源能量调节的能源约束觅食任务中,递归与脉冲智能体表现出适应性行为。冻结基准测试显示,学习后的智能体优于随机与启发式基线;带状态追踪的递归策略表现最佳,脉冲变体呈现压力相关差异。早期内部动态对后期完全安全高效行为具有预测能力,最大ROC-AUC达0.802。降维后主成分子空间仍保留行为相关信号。特征族控制分析表明,预测信号分布在状态追踪、策略头、内部动力学、观测值及反稳态变量中;即使移除显式能量特征,低能状态仍可强解码。评估阶段对时间状态、感官信息、运行条件及反稳态机制的扰动会改变行为与内部预测。种子平衡事件探测显示,未来接触、成功摄食与威胁事件存在较弱但可测量的信息,同时低能状态解码强烈。这一模式可解释为预测性反稳态组织的计算类比:分布式控制机制具备预测性、能量敏感性、动作相关性且部分因果参与,不主张生物学验证或离散符号类别。

原文摘要 · Abstract (English)

Adaptive behavior under partial observability depends on internal organization that carries information beyond the current observation. Drawing on Barrett and Miller's account of categorization as predictive, compressive, functionally organized, and allostatically constrained, we test whether recurrent and spiking agents develop internal states with corresponding computational properties. Agents operate in an energy-constrained foraging task requiring resource acquisition, threat avoidance, contact-dependent consumption, and regulation of an internal energy variable. In a frozen benchmark, learned agents outperform random and heuristic baselines; the trace-augmented recurrent policy is strongest overall, while spiking variants show stress-specific differences. Early internal dynamics predict later full-safe-efficient success above permutation baseline, reaching a maximum ROC-AUC of 0.802. Reduced PCA subspaces retain behaviorally relevant information. Feature-family controls show that predictive signal is distributed across trace, policy-head, internal-dynamics, observation, and allostatic variables, and low-energy state remains strongly decodable after explicit energy-related features are removed. Evaluation-time perturbations to temporal state, sensory information, operating conditions, and allostatic mechanisms alter behavior and/or internal prediction. Seed-balanced event probes show weaker but measurable information about future contact, successful consumption, and threat events, alongside strong low-energy decoding. We interpret this pattern as a computational analogue of predictive allostatic organization: distributed control regimes that are predictive, energy-sensitive, action-relevant, and partly causally involved, without claiming biological validation or discrete symbolic categories.

强化学习神经动力学预测控制能量约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。