针对不同场景动态调整先验知识使用策略,提升强化学习部署灵活性。
Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors

- 根据实际部署情况诊断先验有效性,动态调节其影响
- 实验显示同一先验可能在不同场景中产生助益或损害效果
- 适合需要自适应调整的复杂系统部署,如机器人与大模型后训练
在线强化学习代理越来越依赖离线获取的知识以实现实用效率。这一范式最初源于离线到在线的强化学习,现已扩展至基础模型后训练和具身智能领域,先验知识类型也从离线数据集和预训练策略,拓展到多模态基础模型、生成式世界模型等多样化知识源。离线先验已成为深度强化学习研发与部署的核心。然而,这种依赖带来了当前基准驱动范式无法解决的挑战:先验有效性在不同部署中存在差异,并在训练过程中发生漂移,因此不存在普适最优的管理方法,基准排名对实际部署指导有限。我们主张,领域应转向诊断驱动的张力管理,即通过部署特定证据来指导学习者在整个训练过程中如何与先验互动,从而实现灵活且自适应的部署。我们通过一个框架揭示了先验通过三种功能角色重塑在线优化的过程,辅以控制实验展示帮助与损害效果的反转,跨领域的实证证据(涵盖基础模型后训练至具身智能),以及对五个核心反驳观点的回应,支持该立场。
原文摘要 · Abstract (English)
Online reinforcement learning (RL) agents increasingly depend on knowledge acquired offline to achieve practical efficiency. Originally studied in offline-to-online RL, this paradigm now spans foundation model post-training and embodied intelligence, with prior types expanding from offline datasets and pre-trained policies to increasingly diverse knowledge sources such as multimodal foundation models and generative world models. Offline priors have become central to how deep RL is developed and deployed. However, this reliance introduces a challenge that the prevailing benchmark-driven paradigm cannot resolve: because prior validity varies across deployments and shifts during training, no single approach to managing it is universally optimal, and benchmark rankings offer limited guidance for real-world deployments. Rather than pursuing universal solutions, we argue that the field should shift to diagnosis-driven tension management, in which deployment-specific evidence guides how the learner relates to its priors throughout training, enabling both flexible and adaptive deployment. We support this position with a framework characterizing how priors reshape online optimization through three functional roles, controlled experiments demonstrating help-or-hurt reversals, cross-domain evidence from foundation model post-training to embodied intelligence, and engagement with five substantive counterarguments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。