arXiv:2604.07016cs.LG2026-04

用预测结果构建通用状态表示,实现技能跨任务复用。

Predictive Representations for Skill Transfer in Reinforcement Learning

  • 基于环境结果预测构建任务无关的状态抽象
  • 在未见任务中加速学习,无需预处理
  • 适合需要快速适应新任务的强化学习场景

强化学习规模化的核心挑战在于行为泛化。若无法迁移已学知识,智能体将每次任务都从零开始学习。本文提出基于状态抽象的迁移新范式,通过环境的、与任务无关的紧凑观测(结果)构建结果预测状态表示(OPSRs),即以结果预测为基础的中心化、任务无关状态抽象。理论与实证均表明其具备最优但有限的迁移能力;为此引入基于OPSR的技能(即基于选项的抽象动作),利用状态抽象实现跨任务复用。多组实验中,从示范中学习OPSR技能后,在全新未见任务中显著加快学习速度,且无需任何预处理。我们认为该框架为强化学习中的迁移,特别是状态与动作抽象结合的迁移,提供了有前景的方向。

原文摘要 · Abstract (English)

A key challenge in scaling up Reinforcement Learning is generalizing learned behaviour. Without the ability to carry forward acquired knowledge an agent is doomed to learn each task from scratch. In this paper we develop a new formalism for transfer by virtue of state abstraction. Based on task-independent, compact observations (outcomes) of the environment, we introduce Outcome-Predictive State Representations (OPSRs), agent-centered and task-independent abstractions that are made up of predictions of outcomes. We show formally and empirically that they have the potential for optimal but limited transfer, then overcome this trade-off by introducing OPSR-based skills, i.e. abstract actions (based on options) that can be reused between tasks as a result of state abstraction. In a series of empirical studies, we learn OPSR-based skills from demonstrations and show how they speed up learning considerably in entirely new and unseen tasks without any pre-processing. We believe that the framework introduced in this work is a promising step towards transfer in RL in general, and towards transfer through combining state and action abstraction specifically.

强化学习技能迁移状态抽象预测表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。