arXiv:2605.20272cs.LGcs.AI2026-05

小抽象状态空间让强化学习模型跨尺度泛化

Smaller Abstract State Spaces Enable Cross-Scale Generalization in Reinforcement Learning

  • 通过改进状态抽象框架,构建更紧凑的抽象状态空间
  • 理论证明缩小抽象状态空间可提升分布外泛化性能
  • 适合研究跨任务泛化与智能体架构设计的研究者

尽管人类能将抽象概念推广到更复杂或更大的任务中,但实现强化学习(RL)系统具备此类能力仍具挑战。本文首次提出一种理论模型,说明如何在RL代理中实现分布外(OOD)泛化。方法基于部分可观测马尔可夫决策过程(POMDP),假设智能体通过抽象函数判断哪些经验可视为等价,哪些必须区分。首先将现有状态抽象框架和证明技术扩展至POMDP;其次提出一种成功者加权模型简化方法,允许压缩至比以往定义更小的抽象状态空间。推导出代理在分布外测试中的性能上限,明确界定实现泛化的条件。该边界将性能损失分解为近似误差与估计误差,揭示减小抽象状态空间大小能提升测试表现与分布外泛化能力。分析表明,约束代理在有限小规模抽象状态空间中运行是实现复杂任务泛化的必要条件。结果推动对可跨任务复杂度扩展的强化学习架构的进一步研究。

原文摘要 · Abstract (English)

While humans readily generalize abstract concepts to more complex or larger tasks, building Reinforcement Learning (RL) systems with this ability remains elusive. Here, we present the first theoretical model of how such Out-of-Distribution (OOD) generalization can be achieved in RL agents. Our approach considers Partially Observable Markov Decision Processes (POMDPs) and assumes that an intelligent agent uses an abstraction function to determine which experiences can be treated as equivalent and which must be distinguished. First, we extend the existing state abstraction framework and proof techniques to POMDPs. Then, we define a successor-weighted model reduction, a model reduction variant that enables compression into smaller abstract spaces than prior definitions allow. We derive a bound on the agent's OOD test performance, thereby defining the conditions under which OOD generalization is achievable. This bound decomposes an agent's performance loss into approximation and estimation errors, revealing how reducing an agent's abstract state space size improves test performance and OOD generalization. Our analysis suggests that constraining an agent to operate over a small, finite set of abstract states is necessary for achieving generalization to more complex tasks. Our results motivate further research into learning RL architectures that scale across tasks of varying complexity levels.

强化学习状态抽象泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。