提出自预测框架,让多智能体在动态环境中实现更高阶的协作与推理。
Embedded Universal Predictive Intelligence: a coherent framework for multi-agent learning
- 用自预测机制让智能体同时预测环境输入和自身行为,打破传统分离假设。
- 在多智能体中实现无限阶心智理论,达成经典方法无法达到的合作效果。
- 基于AIXI理论扩展,为理想化嵌入式智能体提供统一学习框架,适合研究者参考。
标准无模型强化学习假设环境动态平稳且智能体与环境解耦,导致多智能体场景中因其他智能体学习引发的非平稳性难以处理。为此,我们基于通用人工智能(AIXI)理论,提出一个以自预测为核心的前瞻性学习与嵌入式智能体框架:贝叶斯强化学习智能体需同时预测未来感知输入和自身行动,并解决对自身作为世界一部分的认知不确定性。在多智能体设置中,该框架使智能体能够推理其他运行相似算法的智能体,产生新的博弈论解概念和经典解耦智能体无法实现的新型合作形式。进一步扩展AIXI理论,研究从Solomonoff先验出发的理想嵌入式智能体,证明其可形成一致的相互预测并实现无限阶心智理论,或将成为嵌入式多智能体学习的黄金标准。
原文摘要 · Abstract (English)
The standard theory of model-free reinforcement learning assumes that the environment dynamics are stationary and that agents are decoupled from their environment, such that policies are treated as being separate from the world they inhabit. This leads to theoretical challenges in the multi-agent setting where the non-stationarity induced by the learning of other agents demands prospective learning based on prediction models. To accurately model other agents, an agent must account for the fact that those other agents are, in turn, forming beliefs about it to predict its future behavior, motivating agents to model themselves as part of the environment. Here, building upon foundational work on universal artificial intelligence (AIXI), we introduce a mathematical framework for prospective learning and embedded agency centered on self-prediction, where Bayesian RL agents predict both future perceptual inputs and their own actions, and must therefore resolve epistemic uncertainty about themselves as part of the universe they inhabit. We show that in multi-agent settings, self-prediction enables agents to reason about others running similar algorithms, leading to new game-theoretic solution concepts and novel forms of cooperation unattainable by classical decoupled agents. Moreover, we extend the theory of AIXI, and study universally intelligent embedded agents which start from a Solomonoff prior. We show that these idealized agents can form consistent mutual predictions and achieve infinite-order theory of mind, potentially setting a gold standard for embedded multi-agent learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。