arXiv:2607.16177cs.LGmath.OC2026-07

用物理规律提升强化学习效率,让机器人更快学会控制复杂系统。

Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems

论文配图:Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems
图 1 · 摘自论文原文
  • 结合物理模型与自动微分,用短时梯度优化策略
  • 只需少量环境交互,就比现有方法更优且稳定
  • 适合高维、参数多的动态系统,无需降维或分拆任务

强化学习(RL)在非线性复杂动力系统中展现出良好的反馈控制潜力,但其样本效率低,需大量环境交互才能获得最优策略。因此,实际应用常受限于稀疏传感器与执行器,尤其在高维空间中面临探索-利用困境带来的维度灾难。本文提出一种新型物理增强强化学习(PEARL)框架,针对高维参数化动力系统,利用系统动力学的可微特性,设计了基于演员-伴随的算法:通过自动微分计算短期策略梯度,并借助神经网络近似未来回报的伴随敏感度,显著减少环境交互次数,同时缓解长期梯度不稳定性。在两个具有挑战性的非定常流场参数化导航任务中验证表明,PEARL(i)有效利用可微环境,在性能上超越当前最先进RL算法;(ii)具备高样本效率,得益于物理引导的学习机制;(iii)可在多种场景间良好泛化,对参数化系统至关重要;(iv)支持在高维状态与动作空间中直接扩展RL,无需低维状态表示或多智能体策略。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has recently emerged as a promising feedback control strategy for nonlinear and complex dynamical systems. However, RL algorithms are sample inefficient and require a large number of interaction with the environment to synthesize optimal control strategies. Consequently, applications of RL are typically limited to sparse sensors and actuators due to the curse of dimensionality entailed by the exploration-exploitation dilemma in high-dimensional spaces. In this work, we bridge RL and traditional optimal control for dynamical system with a novel Physics-EnhAnced Reinforcement Learning (PEARL) paradigm tailored to the control of high-dimensional and parametric dynamical systems, exploiting the differentibility of their dynamics. Specifically, PEARL employs an actor-adjoint algorithm that leverages automatic differentiation to compute policy gradients over short horizons and adjoint-based sensitivities of future returns approximated via neural networks, significantly reducing the number of environment interactions, while mitigating long-term gradient instabilities. Through two challenging parametric navigation problems in unsteady flows, we show that PEARL (i) effectively exploits differentiable environments to outperform state-of-the-art RL algorithms, (ii) is sample efficient, thanks to the physics-guided policy learning, (iii) generalizes across multiple scenarios, which is crucial when dealing with parametric systems, and (iv) enables scaling RL to high-dimensional state and action spaces, without requiring low-dimensional state representations or multi-agent strategies.

强化学习物理模型高效控制动力系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。