arXiv:2412.13106cs.LG2024-12被引 4

提出新方法,用最少在线交互提升离线强化学习性能。

Active Reinforcement Learning Strategies for Offline Policy Improvement

  • 基于已有离线数据主动选择最有价值的在线探索轨迹。
  • 在多个环境上减少75%的额外在线交互,性能优于基线。
  • 适合数据受限或成本高的强化学习场景,如医疗试验与复杂导航。

具备序列决策能力的学习代理需持续解决探索与利用的权衡问题。然而,在线与环境交互可能代价高昂且受约束,如交互预算有限或状态空间某些区域无法探索,例如医学试验候选人筛选和复杂导航环境训练。这要求研究能通过重用先前由未知行为策略收集的离线数据,以最小化额外经验轨迹获取的主动强化学习策略。本文提出一种主动强化学习方法,可有效选择补充现有离线数据的轨迹。大量实验表明,该方法在Gym-MuJoCo运动、Maze2d、AntMaze、CARLA和IsaacSimGo1等多个连续控制环境中,相比竞争基线,最多减少75%的额外在线交互。据我们所知,这是首个在序列决策与强化学习背景下解决主动学习问题的工作。

原文摘要 · Abstract (English)

Learning agents that excel at sequential decision-making tasks must continuously resolve the problem of exploration and exploitation for optimal learning. However, such interactions with the environment online might be prohibitively expensive and may involve some constraints, such as a limited budget for agent-environment interactions and restricted exploration in certain regions of the state space. Examples include selecting candidates for medical trials and training agents in complex navigation environments. This problem necessitates the study of active reinforcement learning strategies that collect minimal additional experience trajectories by reusing existing offline data previously collected by some unknown behavior policy. In this work, we propose an active reinforcement learning method capable of collecting trajectories that can augment existing offline data. With extensive experimentation, we demonstrate that our proposed method reduces additional online interaction with the environment by up to 75% over competitive baselines across various continuous control environments such as Gym-MuJoCo locomotion environments as well as Maze2d, AntMaze, CARLA and IsaacSimGo1. To the best of our knowledge, this is the first work that addresses the active learning problem in the context of sequential decision-making and reinforcement learning.

强化学习主动学习离线策略优化高效探索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。