arXiv:2603.28971eess.SYcs.LG2026-03

HAC通过哈密顿量直接优化策略,避免值函数误差,提升样本效率和鲁棒性。

A Pontryagin Method of Model-based Reinforcement Learning via Hamiltonian Actor-Critic

  • 基于庞特里亚金原理,用哈密顿量替代传统值函数进行策略优化
  • 在连续控制任务中优于基线方法,收敛更快且对分布外数据更鲁棒
  • 特别适合数据有限的离线强化学习场景,样本效率高

基于模型的强化学习(MBRL)通过利用学习到的动力学模型来提升策略优化的样本效率。然而,如演员-评论家这类方法的有效性常受累积模型误差影响,导致长时程价值估计退化。现有方法如基于模型的价值扩展(MVE)虽通过多步回放部分缓解此问题,但仍对回放时长选择敏感且受残余模型偏差影响。受庞特里亚金最大值原理(PMP)启发,本文提出哈密顿演员-评论家(HAC),一种基于模型的方法,通过在学习到的动力学和奖励上直接优化哈密顿量,避免显式值函数学习,适用于确定性系统。通过消除值函数近似,HAC降低了对模型误差的敏感性,并具备收敛性保证。在连续控制基准测试中,无论在线还是离线强化学习设置下,实验表明HAC在控制性能、收敛速度及对分布偏移(包括分布外,OOD)的鲁棒性方面均优于模型无关与基于MVE的基线方法。在数据受限的离线设置中,HAC达到或超越当前最先进水平,凸显其强大的样本效率。

原文摘要 · Abstract (English)

Model-based reinforcement learning (MBRL) improves sample efficiency by leveraging learned dynamics models for policy optimization. However, the effectiveness of methods such as actor-critic is often limited by compounding model errors, which degrade long-horizon value estimation. Existing approaches, such as Model-Based Value Expansion (MVE), partially mitigate this issue through multi-step rollouts, but remain sensitive to rollout horizon selection and residual model bias. Motivated by the Pontryagin Maximum Principle (PMP), we propose Hamiltonian Actor-Critic (HAC), a model-based approach that eliminates explicit value function learning by directly optimizing a Hamiltonian defined over the learned dynamics and reward for deterministic systems. By avoiding value approximation, HAC reduces sensitivity to model errors while admitting convergence guarantees. Extensive experiments on continuous control benchmarks, in both online and offline RL settings, demonstrate that HAC outperforms model-free and MVE-based baselines in control performance, convergence speed, and robustness to distributional shift, including out-of-distribution (OOD) scenarios. In offline settings with limited data, HAC matches or exceeds state-of-the-art methods, highlighting its strong sample efficiency.

强化学习模型预测哈密顿优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。