arXiv:2607.20708cs.LGq-bio.NC2026-07

发现主动推理智能体中因果涌现的根源在慢速全局隐变量,而非训练程度。

Perspective Latents as an Architectural Condition for Causal Emergence in Active Inference Agents

论文配图:Perspective Latents as an Architectural Condition for Causal Emergence in Active Inference Agents
图 1 · 摘自论文原文
  • 用快/慢双隐变量架构分离感知与全局状态,研究因果涌现机制。
  • 因果涌现量Φ_r集中在慢速隐变量g上,且随训练下降而非上升。
  • 适合关注意识理论与主动推理架构的交叉研究者阅读。

近期研究通过整合信息分解衡量强化学习智能体的因果涌现,发现Φ_r随训练增长并追踪奖励提升。对于主动推理框架,这引发了一个问题:无奖励的预测性组织如何与这类信息论指标关联?本文在一个主动推理智能体中测试该问题,其架构将快速感知隐变量z与慢速全局隐变量g分离,其中g由预测误差驱动且与策略梯度结构解耦。在无奖励的环境切换协议下,Φ_r集中于g;其总幅值主要由架构决定,随训练减少。学习的实质性影响仅在原子-组合层面显现:解耦使Φ_r符号由负转正,并在环境变化下保持不变,而向下因果则承载环境依赖的调整。结果表明,g是主动推理智能体中Φ_r相关时间组织的架构核心,反对将标量Φ_r直接视为学习整合性的指标。

原文摘要 · Abstract (English)

A recent line of work measures causal emergence in reinforcement learning agents through Integrated Information Decomposition, reporting that $Φ_r$ grows with training and tracks reward improvement. For active inference, this raises the question of how reward-free predictive organization relates to such information-theoretic signatures. I test this within an active inference agent whose architecture separates a fast perception latent $z$ from a slow global latent $g$, where $g$ is driven by prediction error and structurally decoupled from policy gradients. In a reward-free environmental regime-switching protocol, $Φ_r$ concentrates in $g$; its aggregate magnitude is largely architectural and decreases with training. The substantive effect of learning becomes legible only at the atom-compositional level: decoupling flips sign from negative to positive and becomes regime-invariant under environmental change, while downward causation carries the regime-dependent adjustment. These results identify $g$ as the architectural locus of $Φ_r$-relevant temporal organization in an active inference agent, and argue against reading scalar $Φ_r$ as a direct index of learned integration.

因果涌现主动推理隐变量架构分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。