发现主动推理智能体中因果涌现的根源在慢速全局隐变量,而非训练程度。
Perspective Latents as an Architectural Condition for Causal Emergence in Active Inference Agents

- 用快/慢双隐变量架构分离感知与全局状态,研究因果涌现机制。
- 因果涌现量Φ_r集中在慢速隐变量g上,且随训练下降而非上升。
- 适合关注意识理论与主动推理架构的交叉研究者阅读。
近期研究通过整合信息分解衡量强化学习智能体的因果涌现,发现Φ_r随训练增长并追踪奖励提升。对于主动推理框架,这引发了一个问题:无奖励的预测性组织如何与这类信息论指标关联?本文在一个主动推理智能体中测试该问题,其架构将快速感知隐变量z与慢速全局隐变量g分离,其中g由预测误差驱动且与策略梯度结构解耦。在无奖励的环境切换协议下,Φ_r集中于g;其总幅值主要由架构决定,随训练减少。学习的实质性影响仅在原子-组合层面显现:解耦使Φ_r符号由负转正,并在环境变化下保持不变,而向下因果则承载环境依赖的调整。结果表明,g是主动推理智能体中Φ_r相关时间组织的架构核心,反对将标量Φ_r直接视为学习整合性的指标。
原文摘要 · Abstract (English)
A recent line of work measures causal emergence in reinforcement learning agents through Integrated Information Decomposition, reporting that $Φ_r$ grows with training and tracks reward improvement. For active inference, this raises the question of how reward-free predictive organization relates to such information-theoretic signatures. I test this within an active inference agent whose architecture separates a fast perception latent $z$ from a slow global latent $g$, where $g$ is driven by prediction error and structurally decoupled from policy gradients. In a reward-free environmental regime-switching protocol, $Φ_r$ concentrates in $g$; its aggregate magnitude is largely architectural and decreases with training. The substantive effect of learning becomes legible only at the atom-compositional level: decoupling flips sign from negative to positive and becomes regime-invariant under environmental change, while downward causation carries the regime-dependent adjustment. These results identify $g$ as the architectural locus of $Φ_r$-relevant temporal organization in an active inference agent, and argue against reading scalar $Φ_r$ as a direct index of learned integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。