用多项式能量模型构建可精确计算熵的策略,解决多模态决策难题
MePoly: Max Entropy Polynomial Policy Optimization
- 基于多项式能量模型设计新策略参数化方式
- 在多个基准上性能优于现有方法,能捕捉复杂非凸流形
- 适合需要高精度熵优化的强化学习与模仿学习场景
随机最优控制为复杂决策问题提供了统一的数学框架,涵盖最大熵强化学习(RL)和模仿学习(IL)等范式。然而,传统参数化策略常难以表示解的多模态特性。尽管基于扩散的方法旨在恢复多模态性,但其缺乏显式的概率密度,使策略梯度优化变得困难。为此,我们提出MePoly,一种基于多项式能量模型的新策略参数化方法。MePoly提供显式且可处理的概率密度,支持精确熵最大化。理论上,我们基于经典矩问题,利用任意分布的通用逼近能力。实验上,我们证明MePoly能有效捕捉复杂的非凸流形,并在多种基准测试中表现优于基线方法。
原文摘要 · Abstract (English)
Stochastic Optimal Control provides a unified mathematical framework for solving complex decision-making problems, encompassing paradigms such as maximum entropy reinforcement learning(RL) and imitation learning(IL). However, conventional parametric policies often struggle to represent the multi-modality of the solutions. Though diffusion-based policies are aimed at recovering the multi-modality, they lack an explicit probability density, which complicates policy-gradient optimization. To bridge this gap, we propose MePoly, a novel policy parameterization based on polynomial energy-based models. MePoly provides an explicit, tractable probability density, enabling exact entropy maximization. Theoretically, we ground our method in the classical moment problem, leveraging the universal approximation capabilities for arbitrary distributions. Empirically, we demonstrate that MePoly effectively captures complex non-convex manifolds and outperforms baselines in performance across diverse benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。