arXiv:2412.01245cs.LGcs.AI2024-12被引 3

提出更简单的生成式策略算法,提升连续动作强化学习效果。

Revisiting Generative Policies: A Simpler Reinforcement Learning Algorithmic Perspective

  • 用优势加权回归简化训练目标,降低算法复杂度。
  • 新方法在多个离线强化学习数据集上达到顶尖性能。
  • 适合关注生成模型与强化学习结合的研究者参考。

生成模型(尤其是扩散模型)在多模态数据密度估计方面取得显著成功,引发强化学习领域对连续动作空间策略建模的广泛关注。然而,现有方法在训练方案和优化目标上差异较大,部分仅适用于扩散模型。本文系统比较并分析了多种生成式策略的训练与部署技术,识别并验证了有效的算法设计。具体地,我们重新审视现有训练目标,将其分为两类,并对应提出两种更简洁的方法:第一种是生成模型策略优化(GMPO),采用原生的优势加权回归作为训练目标,显著简化此前复杂方法;第二种是生成模型策略梯度(GMPG),提供一种数值稳定的原生策略梯度实现。我们构建了一个标准化实验框架GenerativeRL。实验表明,所提方法在多个离线强化学习数据集上达到当前最优表现,为生成式策略的训练与部署提供了统一且实用的指导。

原文摘要 · Abstract (English)

Generative models, particularly diffusion models, have achieved remarkable success in density estimation for multimodal data, drawing significant interest from the reinforcement learning (RL) community, especially in policy modeling in continuous action spaces. However, existing works exhibit significant variations in training schemes and RL optimization objectives, and some methods are only applicable to diffusion models. In this study, we compare and analyze various generative policy training and deployment techniques, identifying and validating effective designs for generative policy algorithms. Specifically, we revisit existing training objectives and classify them into two categories, each linked to a simpler approach. The first approach, Generative Model Policy Optimization (GMPO), employs a native advantage-weighted regression formulation as the training objective, which is significantly simpler than previous methods. The second approach, Generative Model Policy Gradient (GMPG), offers a numerically stable implementation of the native policy gradient method. We introduce a standardized experimental framework named GenerativeRL. Our experiments demonstrate that the proposed methods achieve state-of-the-art performance on various offline-RL datasets, offering a unified and practical guideline for training and deploying generative policies.

强化学习生成模型策略优化扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。