arXiv:2411.09722cs.LGcs.AI2024-11被引 1

通过迭代优化,让离线强化学习持续收集安全且多样化的数据以提升策略。

Iterative Batch Reinforcement Learning via Safe Diversified Model-based Policy Search

  • 基于集成模型的策略搜索,结合安全与多样性约束。
  • 在不偏离已有数据分布的前提下,逐步提升策略性能。
  • 适合工业控制等高风险场景中需持续迭代优化的系统。

批量强化学习允许在训练期间无需与环境直接交互,仅依赖预先收集的数据集进行策略学习,因此特别适用于高风险、高成本的应用场景,如工业控制。当前学习的策略通常受限于批次数据中的行为模式。在实际部署中,策略运行会生成新数据,这些数据可被添加到原有记录中,形成迭代学习闭环。本文提出一种基于集成模型的迭代批量强化学习方法,引入安全约束和关键的多样性准则,引导策略在部署阶段主动采集高效且信息丰富的数据,实现策略的持续改进,同时保持在已收集数据的支持范围内。

原文摘要 · Abstract (English)

Batch reinforcement learning enables policy learning without direct interaction with the environment during training, relying exclusively on previously collected sets of interactions. This approach is, therefore, well-suited for high-risk and cost-intensive applications, such as industrial control. Learned policies are commonly restricted to act in a similar fashion as observed in the batch. In a real-world scenario, learned policies are deployed in the industrial system, inevitably leading to the collection of new data that can subsequently be added to the existing recording. The process of learning and deployment can thus take place multiple times throughout the lifespan of a system. In this work, we propose to exploit this iterative nature of applying offline reinforcement learning to guide learned policies towards efficient and informative data collection during deployment, leading to continuous improvement of learned policies while remaining within the support of collected data. We present an algorithmic methodology for iterative batch reinforcement learning based on ensemble-based model-based policy search, augmented with safety and, importantly, a diversity criterion.

强化学习离线学习工业控制迭代优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。