arXiv:2502.07937cs.LGstat.ML2025-02被引 6

A3RL通过动态选择在线与离线数据,提升强化学习采样效率。

Active Advantage-Aligned Online Reinforcement Learning with Offline Data

  • 基于策略需求动态筛选在线与离线数据,实现主动采样。
  • 在多种环境中优于依赖离线数据的在线强化学习方法。
  • 适合需要高效迭代且数据质量不一的强化学习场景。

在线强化学习(RL)通过与环境直接交互来优化策略,但面临样本效率低的问题。离线RL利用大量预收集数据学习策略,却因数据覆盖有限而表现不佳。近期研究尝试结合两者优势,但存在灾难性遗忘、对数据质量敏感及数据利用效率低等问题。为此,我们提出A3RL,引入一种新型置信度感知的主动优势对齐(A3)采样策略,动态从在线和离线数据中优先选择符合策略演化需求的数据,以优化策略改进。我们还提供了该主动采样策略有效性的理论分析,并进行了多样化的实验与消融研究,结果表明,本方法在性能上优于其他利用离线数据的在线强化学习技术。

原文摘要 · Abstract (English)

Online reinforcement learning (RL) enhances policies through direct interactions with the environment, but faces challenges related to sample efficiency. In contrast, offline RL leverages extensive pre-collected data to learn policies, but often produces suboptimal results due to limited data coverage. Recent efforts integrate offline and online RL in order to harness the advantages of both approaches. However, effectively combining online and offline RL remains challenging due to issues that include catastrophic forgetting, lack of robustness to data quality and limited sample efficiency in data utilization. In an effort to address these challenges, we introduce A3RL, which incorporates a novel confidence aware Active Advantage Aligned (A3) sampling strategy that dynamically prioritizes data aligned with the policy's evolving needs from both online and offline sources, optimizing policy improvement. Moreover, we provide theoretical insights into the effectiveness of our active sampling strategy and conduct diverse empirical experiments and ablation studies, demonstrating that our method outperforms competing online RL techniques that leverage offline data.

强化学习在线学习数据融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。