通用智能体需记忆关键信息以在多场景中近优决策。
What Must Generalist Agents Remember?

- 当环境观测重叠但最优动作冲突时,记忆必须不同
- 仅靠当前状态无法近优决策,必须保留领域相关记忆
- 记忆可重建局部动态模型,支持规划与推理
本文建立了一个形式化框架,阐明通用智能体为在多个环境与目标中近优行动,必须在记忆中存储何种信息。研究发现,当两个领域存在共同观测瓶颈但要求不兼容的最优行为时,任何近优策略都必须在该瓶颈处诱导不同的记忆分布。这一结果推出一个分离定理:足够成功的智能体不能仅依赖当前状态观测,而必须在记忆中保留领域相关的信息。进一步证明,若智能体记忆中包含足够信息以估计相关目标的价值,则该记忆可用于近似重构局部转移动态。综合来看,这些结果将记忆定义为支持领域辨识、转移模型重建和规划的核心载体,对通用智能体的设计具有根本性指导意义。
原文摘要 · Abstract (English)
This paper develops a formal account of what generalist agents must store in memory in order to act near-optimally across multiple environments and goals. It shows that when two domains share an observational bottleneck but require incompatible optimal actions, any uniformly near-optimal policy must induce distinct memory distributions at that bottleneck. The result yields a separation theorem: sufficiently successful agents cannot rely only on current state observations, but must preserve domain-relevant information in memory. The paper further shows that if an agent's memory contains enough information to estimate values for related goals, then that memory can be used to approximately reconstruct the agent's local transition dynamics. Together, these results characterize memory as the substrate that supports domain disambiguation, transition-model reconstruction, and planning for generalist agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。