arXiv:2501.02774cs.LG2025-01

提出FLEXplore模型,提升复杂参数化动作环境下的强化学习效率与探索能力。

Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes

  • 基于动态建模与路径积分控制,构建灵活的参数化动作策略。
  • 在多个标准基准上实现更快收敛与更优长期性能,超越现有基线。
  • 适合需要高效探索与高维动作空间建模的强化学习任务。

混合动作模型在强化学习中被广泛认为是有效的方法。当前主流方法是在参数化动作马尔可夫决策过程(PAMDP)下训练智能体,虽在特定环境中表现良好,但在复杂PAMDP中学习效率显著下降,或在原始空间与隐空间转换时丢失关键信息。为提升智能体的学习效率与最终性能,我们提出一种基于模型的强化学习算法FLEXplore。该方法学习参数化动作条件下的动态模型,并采用改进的模型预测路径积分控制。不同于传统基于模型的强化学习算法,我们精心设计动态损失函数与奖励平滑过程,以学习一个宽松但灵活的模型。此外,通过变分下界最大化状态与混合动作之间的互信息,提升探索有效性。理论上证明,在给定Lipschitz条件下,FLEXplore可通过Wasserstein度量降低滚动轨迹的遗憾。在多个标准基准上的实验结果表明,相较于其他基线,FLEXplore具有卓越的学习效率与最终性能。

原文摘要 · Abstract (English)

Hybrid action models are widely considered an effective approach to reinforcement learning (RL) modeling. The current mainstream method is to train agents under Parameterized Action Markov Decision Processes (PAMDPs), which performs well in specific environments. Unfortunately, these models either exhibit drastic low learning efficiency in complex PAMDPs or lose crucial information in the conversion between raw space and latent space. To enhance the learning efficiency and asymptotic performance of the agent, we propose a model-based RL (MBRL) algorithm, FLEXplore. FLEXplore learns a parameterized-action-conditioned dynamics model and employs a modified Model Predictive Path Integral control. Unlike conventional MBRL algorithms, we carefully design the dynamics loss function and reward smoothing process to learn a loose yet flexible model. Additionally, we use the variational lower bound to maximize the mutual information between the state and the hybrid action, enhancing the exploration effectiveness of the agent. We theoretically demonstrate that FLEXplore can reduce the regret of the rollout trajectory through the Wasserstein Metric under given Lipschitz conditions. Our empirical results on several standard benchmarks show that FLEXplore has outstanding learning efficiency and asymptotic performance compared to other baselines.

强化学习模型预测探索优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。