arXiv:2603.08468eess.SYcs.LG2026-03被引 1

用拉格朗日神经网络提升强化学习模型的物理一致性与训练效率。

Integrating Lagrangian Neural Networks into the Dyna Framework for Reinforcement Learning

  • 在Dyna框架中引入拉格朗日神经网络,强制模型遵循物理规律。
  • 基于状态估计的优化方法比随机梯度法收敛更快。
  • 适合关注物理引导强化学习与高效训练的科研人员。

基于模型的强化学习(MBRL)具有样本高效性,但其性能依赖于动态模型的准确性。现有方法多采用黑箱建模,难以满足物理规律约束,在分布外数据上预测偏差大。本文将拉格朗日神经网络(LNN)融入基于Dyna的MBRL框架,通过强制学习过程遵循拉格朗日结构,提升模型泛化能力。同时,对比了基于随机梯度和状态估计的两种优化方法,结果表明状态估计法在神经网络训练中收敛速度更快。仿真结果验证了该框架在动态建模中的有效性。

原文摘要 · Abstract (English)

Model-based reinforcement learning (MBRL) is sample-efficient but depends on the accuracy of the learned dynamics, which are often modeled using black-box methods that do not adhere to physical laws. Those methods tend to produce inaccurate predictions when presented with data that differ from the original training set. In this work, we employ Lagrangian neural networks (LNNs), which enforce an underlying Lagrangian structure to train the model within a Dyna-based MBRL framework. Furthermore, we train the LNN using stochastic gradient-based and state-estimation-based optimizers to learn the network's weights. The state-estimation-based method converges faster than the stochastic gradient-based method during neural network training. Simulation results are provided to illustrate the effectiveness of the proposed LNN-based Dyna framework for MBRL.

强化学习物理模型神经网络动态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。