arXiv:2504.02543cs.LG2025-04被引 4

用均值哈密顿量最小化实现不确定系统最优控制

Optimal Control of Probabilistic Dynamics Models via Mean Hamiltonian Minimization

  • 基于后验分布最小化均值哈密顿量,处理模型不确定性
  • 在离线任务中降低试错成本,在线任务表现与基准相当
  • 适用于学习型动态模型,尤其适合大规模神经ODE

在缺乏真实系统动力学精确知识的情况下,非线性连续时间系统的最优控制需谨慎处理认知不确定性。本文将庞特里亚金最大值原理的概率解释转化为学习型概率动力学模型下的最优控制问题。通过在系统动力学的后验分布上最小化均值哈密顿量,提供了一种原则性的不确定性处理方法。我们提出一种多射击数值方法,利用均值哈密顿量最小化,并可扩展至大规模概率动力学模型,包括集成神经微分方程。在在线与离线模型基强化学习任务中的对比表明,该概率哈密顿方法在离线设置中显著降低试错成本,在在线场景中表现具有竞争力。本方法连接了最优控制与强化学习,为学习型不确定系统的控制提供了原则性且实用的框架。

原文摘要 · Abstract (English)

Without exact knowledge of the true system dynamics, optimal control of non-linear continuous-time systems requires careful treatment under epistemic uncertainty. In this work, we translate a probabilistic interpretation of the Pontryagin maximum principle to the challenge of optimal control with learned probabilistic dynamics models. Our framework provides a principled treatment of epistemic uncertainty by minimizing the mean Hamiltonian with respect to a posterior distribution over the system dynamics. We propose a multiple shooting numerical method that leverages mean Hamiltonian minimization and is scalable to large-scale probabilistic dynamics models, including ensemble neural ordinary differential equations. Comparisons against other baselines in online and offline model-based reinforcement learning tasks show that our probabilistic Hamiltonian approach leads to reduced trial costs in offline settings and achieves competitive performance in online scenarios. By bridging optimal control and reinforcement learning, our approach offers a principled and practical framework for controlling uncertain systems with learned dynamics.

最优控制概率模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。