提出计算嵌入视角,揭示智能体永远受限于环境的本质。
The World Is Bigger! A Computationally-Embedded Perspective on the Big World Hypothesis
- 将智能体视为在通用计算机中运行的自动机,天然受环境约束
- 发现深度线性网络比非线性网络更持久保持适应能力
- 适合研究持续学习、智能体与环境关系的学者参考
持续学习常基于‘世界比智能体大’的大世界假说。现有方法通过显式约束智能体来体现这一思想,导致智能体不断适应有限能力,而非收敛。但此类约束往往主观、难融入且阻碍规模扩展。本文提出计算嵌入视角:将智能体视为在通用(形式化)计算机中模拟的自动机,始终处于环境嵌入状态。我们证明该设定等价于在可数无穷状态空间的局部可观测马尔可夫决策过程中的交互。为此提出‘互动性’目标,衡量智能体通过学习新预测持续调整行为的能力。进而设计基于模型的强化学习算法以追求互动性,并构建合成任务评估持续学习性能。结果表明,深度非线性网络难以维持互动性,而深度线性网络随容量增加能保持更高互动性。
原文摘要 · Abstract (English)
Continual learning is often motivated by the idea, known as the big world hypothesis, that "the world is bigger" than the agent. Recent problem formulations capture this idea by explicitly constraining an agent relative to the environment. These constraints lead to solutions in which the agent continually adapts to best use its limited capacity, rather than converging to a fixed solution. However, explicit constraints can be ad hoc, difficult to incorporate, and may limit the effectiveness of scaling up the agent's capacity. In this paper, we characterize a problem setting in which an agent, regardless of its capacity, is constrained by being embedded in the environment. In particular, we introduce a computationally-embedded perspective that represents an embedded agent as an automaton simulated within a universal (formal) computer. Such an automaton is always constrained; we prove that it is equivalent to an agent that interacts with a partially observable Markov decision process over a countably infinite state-space. We propose an objective for this setting, which we call interactivity, that measures an agent's ability to continually adapt its behaviour by learning new predictions. We then develop a model-based reinforcement learning algorithm for interactivity-seeking, and use it to construct a synthetic problem to evaluate continual learning capability. Our results show that deep nonlinear networks struggle to sustain interactivity, whereas deep linear networks sustain higher interactivity as capacity increases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。