arXiv:2607.25207cs.LG2026-07被引 1

解决离线数据与在线环境动态不一致的强化学习问题

A Unified Algorithmic Framework for Hybrid Reinforcement Learning in Tabular MDPs with Shifted Transition Dynamics

  • 提出统一框架,用偏差信息优化离线数据利用
  • 理论证明算法在多种场景下均能实现最优性能
  • 适合研究迁移学习与稳健强化学习的研究者

本文研究了在表格型马尔可夫决策过程(Tabular MDPs)中的混合强化学习设置,其中智能体需结合对目标环境的在线交互和来自源环境的离线数据来学习最优策略。核心挑战在于离线数据可能来自过渡动态已发生偏移的旧环境,导致直接使用历史数据效果不佳。为此,我们提出一个统一算法框架,包含两个算法:MIN-UCB-VI用于最小化累积遗憾,MAX-LCB-VI用于识别最优策略。两者均利用细粒度偏差信息,在一般转移动态偏移下更有效地利用离线数据。我们为该框架提供了理论保证,包括实例相关与无关的遗憾和次优性差距上界。此外,我们建立了匹配的下界以证明方法的最优性,并通过大量实验验证了理论结果。

原文摘要 · Abstract (English)

This paper investigates a hybrid reinforcement learning setting in tabular Markov Decision Processes (MDPs), where an agent aims to learn an optimal policy by combining online interactions with a target environment and offline data from a source environment. A central challenge is that offline data may be collected from outdated environments with shifted transition dynamics, making naive integration of historical data ineffective. To address this, we propose a unified algorithmic framework featuring two algorithms: MIN-UCB-VI for regret minimization and MAX-LCB-VI for best policy identification. Both algorithms leverage fine-grained bias information to more effectively exploit offline data under general transition shifts. We provide theoretical guarantees for our framework, including both instance-dependent and independent upper bounds on regret and sub-optimality gap. Furthermore, we establish matching lower bounds to demonstrate the optimality of our approach and validate our theoretical findings through extensive experiments.

强化学习离线学习动态偏移理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。