提出代理行为的理论框架,区分预测、压缩与赋能机制。
Prediction and Empowerment: A Theory of Agency through Bridge Interfaces
- 将感知与行动建模为跨主体的接口,定义确定性部分可观测环境下的代理行为。
- 证明完美预测需识别隐含状态或通过控制使未来动作决定,仅高赋能不足。
- 强调现代AI设计应区分状态识别、接口优化、任务可控性等目标,适合对齐研究者。
我们研究在确定性物理或模拟世界中部分可观测条件下的代理行为,其中看似随机性源于初始状态、固定规律位及外生噪声的不确定性。将感知与执行建模为分属代理控制参数与环境控制信道状态的桥接接口,通过潜在微观状态的先验分布和多对一观测粗化,构建确定性部分可观测马尔可夫决策过程(POMDP)。在此框架下,我们证明了预测、压缩与赋能之间的分离性:完美预测可通过识别目标相关隐含商集或通过覆盖控制使未来目标动作决定实现;仅高赋能不足以达成。在可细化接口和充足记忆条件下,以动作条件的观测压缩可降低对潜在商集后验不确定性的认知,当精炼需引导世界侧信道条件时,形成目标条件接口赋能。比特串特化模型在信息预算恒定下显式展现权衡:通过识别实现预测需内部容量至少达到相关隐含熵,而覆盖控制则需终端动作容量超过控制商集。对现代人工智能代理而言,结果指向一种设计原则而非必然定理:目标应区分隐藏状态识别、接口精炼、任务相关可控性以及单纯覆盖或干扰控制。人机对齐本质上是接口设计问题,关键桥梁在于人类意图、代理内部状态、外部工具与世界侧信道条件之间。此为工作草稿,欢迎反馈与批评。
原文摘要 · Abstract (English)
We study agency under partial observability in deterministic physical or simulated worlds, where apparent randomness arises from uncertainty over initial conditions, fixed law bits, and unrolled exogenous noise. We model sensing and actuation as bridge interfaces split between agent-controlled parameters and environment-controlled channel state, inducing a deterministic POMDP through a prior over latent microstates and many-to-one observation coarsening. Within this framework, we prove a separation between prediction, compression, and empowerment. Perfect prediction can be achieved either by identifying the hidden quotient relevant to the target family or by overwrite control that makes the future target action-determined; high empowerment alone is insufficient. Under refinable interfaces and sufficient memory, action-conditioned observation-compression progress reduces posterior uncertainty about the latent quotient, and when refinement requires steering world-side channel conditions, this creates target-conditioned interface empowerment. A bit-string specialization with a conserved information budget makes the resulting tradeoff explicit: prediction by identification requires internal capacity at least the relevant latent entropy, whereas overwrite control requires terminal action capacity over the controlled quotient. For modern AI agents, the results suggest a design principle rather than a theorem of inevitability: objectives should distinguish hidden-state identification, interface refinement, task-relevant controllability, and mere overwrite or distractor control. Human--AI alignment is partly an interface-design problem, where the relevant bridge is between human intent, agent internal state, external tools, and world-side channel conditions. This is a working draft: feedback and criticism is most welcome.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。