arXiv:2505.18447cs.LG2025-05ICML被引 3

用悲观原则提升迁移强化学习的安全性与可靠性

Pessimism Principle Can Be Effective: Towards a Framework for Zero-Shot Transfer Reinforcement Learning

  • 基于悲观原则构建目标环境的保守性能估计
  • 提供目标性能下界,确保决策安全可靠
  • 支持多源迁移且避免负迁移,适合高风险场景

迁移强化学习旨在利用相关源域的丰富数据,在目标环境数据有限的情况下获得近最优策略。然而,该方法面临两大挑战:转移策略缺乏性能保证,可能导致不良行为;当涉及多个源域时存在负迁移风险。本文提出一种基于悲观原则的新框架,通过构建并优化目标域性能的保守估计,有效解决上述问题。该框架不仅能提供目标性能的优化下界,保障决策安全性,还表现出随源域质量提升而单调改进的特性,从而避免负迁移。我们构造了两种保守估计形式,严格刻画其有效性,并设计了具有收敛保证的高效分布式算法。该框架为强化学习中的迁移学习提供了理论严谨且实践稳健的解决方案。

原文摘要 · Abstract (English)

Transfer reinforcement learning aims to derive a near-optimal policy for a target environment with limited data by leveraging abundant data from related source domains. However, it faces two key challenges: the lack of performance guarantees for the transferred policy, which can lead to undesired actions, and the risk of negative transfer when multiple source domains are involved. We propose a novel framework based on the pessimism principle, which constructs and optimizes a conservative estimation of the target domain's performance. Our framework effectively addresses the two challenges by providing an optimized lower bound on target performance, ensuring safe and reliable decisions, and by exhibiting monotonic improvement with respect to the quality of the source domains, thereby avoiding negative transfer. We construct two types of conservative estimations, rigorously characterize their effectiveness, and develop efficient distributed algorithms with convergence guarantees. Our framework provides a theoretically sound and practically robust solution for transfer learning in reinforcement learning.

强化学习迁移学习安全决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。