通过注入噪声提升生成策略的探索能力,实现离线到在线强化学习的高效迁移。
Flow Matching with Injected Noise for Offline-to-Online Reinforcement Learning
- 基于流匹配构建策略,注入噪声以扩大动作探索范围。
- 在有限在线预算下,性能显著优于现有方法,提升样本效率。
- 适合需要高效在线适应的复杂任务,如机器人控制与游戏智能体。
生成模型在多个领域表现卓越,推动其作为强化学习中表达能力强的策略应用。尽管在离线强化学习中表现优异,尤其当目标分布明确时,其向在线微调的扩展仍被视为离线预训练的直接延续,未解决关键挑战。本文提出一种新方法——离线到在线强化学习中的注入噪声流匹配(FINO),利用基于流匹配的策略提升离线到在线强化学习的样本效率。FINO通过在策略训练中注入噪声,促进更广泛的动作探索,超越离线数据集所见动作范围。此外,结合熵引导采样机制,在在线微调过程中平衡探索与利用,使策略能动态调整行为。在多种高难度任务上的实验表明,即使在线预算有限,FINO也能持续实现更优性能。
原文摘要 · Abstract (English)
Generative models have recently demonstrated remarkable success across diverse domains, motivating their adoption as expressive policies in reinforcement learning (RL). While they have shown strong performance in offline RL, particularly where the target distribution is well defined, their extension to online fine-tuning has largely been treated as a direct continuation of offline pre-training, leaving key challenges unaddressed. In this paper, we propose Flow Matching with Injected Noise for Offline-to-Online RL (FINO), a novel method that leverages flow matching-based policies to enhance sample efficiency for offline-to-online RL. FINO facilitates effective exploration by injecting noise into policy training, thereby encouraging a broader range of actions beyond those observed in the offline dataset. In addition to exploration-enhanced flow policy training, we combine an entropy-guided sampling mechanism to balance exploration and exploitation, allowing the policy to adapt its behavior throughout online fine-tuning. Experiments across diverse, challenging tasks demonstrate that FINO consistently achieves superior performance under limited online budgets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。