用自由能最小化框架,量化AI系统的代理行为类型。
Active Inference: A method for Phenotyping Agency in AI systems?

- 基于信念与偏好推导行动,通过变分推断构建代理链。
- 在T型迷宫中,用通道容量区分低、中、高代理表型。
- 适合关注AI治理与内在动机建模的研究者。
随着代理型人工智能的兴起,现有概念工具难以有效刻画计算系统中的代理性。本文提出一种满足三重标准的最小化代理定义:意图性(基于信念与欲望的行动)、合理性(世界模型所蕴含的规范一致行动)和可解释性(行动可追溯至内部状态)。我们将其形式化为部分可观测马尔可夫决策过程下的变分框架,其中后验信念、先验偏好与期望自由能最小化共同构成代理行动链。基于经典T型迷宫范式,我们证明赋能(empowerment,即动作与预期观测间的信道容量)可作为操作性度量,通过生成模型的结构扰动,有效区分零、中、高代理表型。最后指出,当代理系统主动探索以消除不确定性时,有效的治理机制必须从外部约束转向对先验偏好的内在调节,从而建立计算表型到AI治理策略的原理性桥梁。
原文摘要 · Abstract (English)
The proliferation of agentic artificial intelligence has outpaced the conceptual tools needed to characterize agency in computational systems. Prevailing definitions mainly rely on autonomy and goal-directedness. Here, we argue for a minimal notion open to principled inspection given three criteria: intentionality as action grounded in beliefs and desires, rationality as normatively coherent action entailed by a world model, and explainability as action causally traceable to internal states; we subsequently instantiate these as a partially observable Markov decision process under a variational framework wherein posterior beliefs, prior preferences, and the minimization of expected free energy jointly constitute an agentic action chain. Using a canonical T-maze paradigm, we evidence how empowerment, formulated as the channel capacity between actions and anticipated observations, serves as an operational metric that distinguishes zero-, intermediate-, and high-agency phenotypes through structural manipulations of the generative model. We conclude by arguing that as agents engage in epistemic foraging to resolve ambiguity, the governance controls that remain effective must shift systematically from external constraints to the internal modulation of prior preferences, offering a principled, variational bridge from computational phenotyping to AI governance strategy
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。