用高阶因果函数构建智能体策略,揭示不确定因果结构的优势
Agent policies from higher-order causal functions
- 将智能体策略映射为高阶过程函数,建立跨领域理论桥梁
- 在去中心化部分可观测马尔可夫决策问题中,不确定因果结构性能更优
- 适合研究因果推理、多智能体系统与量子计算交叉的学者
我们建立了确定性部分可观测马尔可夫决策过程(POMDP)中智能体-状态策略等价类与单输入过程函数(高阶量子操作的经典确定性极限)之间的对应关系。利用这一对应,搭建了人工智能中的智能体-环境交互、物理基础中的因果结构与计算机科学中的逻辑之间的桥梁。我们构造了一个*-自伴范畴PF,支持对策略单步评估及多智能体观测约束的切割与张量积解释。在类型层面,我们将无观测依赖的去中心化POMDP识别为建模不确定因果性的多输入过程函数的自然域。进一步证明,在此类去中心化POMDP上,使用不确定因果结构的策略与固定背景因果结构的策略存在严格性能差距:存在实例显示,前者能获得更高的有限时域回报。
原文摘要 · Abstract (English)
We establish a correspondence between equivalence classes of agent-state policies for deterministic POMDPs and one-input process functions (the classical-deterministic limit of higher-order quantum operations). We use this correspondence to build a bridge between the agent-environment interaction in artificial intelligence, causal structure in the foundations of physics, and logic in computer science. We construct a *-autonomous category PF of types which supports an interpretation of one-step evaluation of policies, and multi-agent observation constraints, into cuts and monoidal products. In terms of types, we develop the correspondence further by identifying observation-independent decentralised POMDPs as the natural domain for the multi-input process functions used to model indefinite causality. We then prove a strict separation between general multi-input process function and definite-ordered process function performance on such dec-POMDPs, by finding an instance for which policies utilizing an indefinite causal structure can achieve greater finite-horizon rewards than policies which are restricted to a fixed background causal structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。