arXiv:2510.01460cs.LGcs.AI2025-10被引 3

提出三阶段框架解释离线到在线强化学习的不一致表现

The Three Regimes of Offline-to-Online Reinforcement Learning

  • 基于稳定-可塑性原则划分三种在线微调阶段
  • 45/63实验结果与框架预测高度一致,仅3次相反
  • 适合研究离线预训练与在线微调设计的学者

离线到在线强化学习(RL)已成为一种实用范式,利用离线数据集进行预训练,并通过在线交互进行微调。然而其经验行为极不稳定:在某一设置中表现良好的在线微调设计,在另一设置中可能完全失效。受稳定性-可塑性原则启发,我们提出一个可解释该不一致性的框架:高效微调必须在保留更强离线先验(预训练策略或离线数据集)效用的同时保持足够可塑性。这一视角识别出三种不同的在线微调阶段,每种需具备不同稳定性特性。通过大规模实证研究验证,63个案例中有45个结果与框架预测高度一致,仅有3个相反。本工作为基于离线数据集与预训练策略相对性能的离线到在线RL设计提供指导框架。

原文摘要 · Abstract (English)

Offline-to-online reinforcement learning (RL) has emerged as a practical paradigm that leverages offline datasets for pretraining and online interactions for fine-tuning. However, its empirical behavior is highly inconsistent: design choices of online fine-tuning that work well in one setting can fail completely in another. Guided by the stability--plasticity principle, we propose a framework that can explain this inconsistency: We argue that efficient fine-tuning must preserve the utility of the stronger offline prior, whether that is the pretrained policy or the offline dataset, while maintaining sufficient plasticity. This perspective identifies three regimes of online fine-tuning, each requiring distinct stability properties. We validate this framework through a large-scale empirical study, finding that the results strongly align with its predictions in 45 out of 63 cases, with only 3 opposite mismatches. This work provides a framework for guiding design choices in offline-to-online RL based on the relative performance of the offline dataset and the pretrained policy.

强化学习离线学习在线微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。