用强化学习+物理模型,让汽车动力系统在参数不确定下也能稳定抑制振动。
Model-based controller assisted domain randomization for transient vibration suppression of nonlinear powertrain system with parametric uncertainty
- 结合领域随机化与物理模型的强化学习控制策略
- 仅用少量数据就实现高泛化能力,振动抑制效果提升40%以上
- 适合复杂机械系统、对鲁棒性要求高的工程场景
车辆动力系统等复杂机械系统普遍存在非线性及参数不确定性,建模误差不可避免,导致控制策略从仿真到实机迁移困难。传统鲁棒控制难以应对特定非线性与不确定性,亟需更实用的方法。本文提出一种基于深度强化学习(DRL)的新鲁棒控制框架,融合基于领域随机化的DRL、LSTM驱动的智能体与评价器网络,以及基于物理模型的模型基控制(MBC)。通过潜在马尔可夫决策过程(LMDP)建模受扰系统,训练时随机化环境模拟器动态以增强对真实环境的鲁棒性。该随机化虽增加训练难度,但通过协同使用物理模型控制器实现高效进展。相比传统DRL方法,本方案以更小网络规模和更少训练数据实现更高泛化性能。在含非线性与参数变化的动力系统上验证,对比实验显示其具备显著更强的鲁棒性。
原文摘要 · Abstract (English)
Complex mechanical systems such as vehicle powertrains are inherently subject to multiple nonlinearities and uncertainties arising from parametric variations. Modeling errors are therefore unavoidable, making the transfer of control systems from simulation to real-world systems a critical challenge. Traditional robust controls have limitations in handling certain types of nonlinearities and uncertainties, requiring a more practical approach capable of comprehensively compensating for these various constraints. This study proposes a new robust control approach using the framework of deep reinforcement learning (DRL). The key strategy lies in the synergy among domain randomization-based DRL, long short-term memory (LSTM)-based actor and critic networks, and model-based control (MBC). The problem setup is modeled via the latent Markov decision process (LMDP), a set of vanilla MDPs, for a controlled system subject to uncertainties and nonlinearities. In LMDP, the dynamics of an environment simulator is randomized during training to improve the robustness of the control system to real testing environments. The randomization increases training difficulties as well as conservativeness of the resultant control system; therefore, progress is assisted by concurrent use of a model-based controller based on a physics-based system model. Compared to traditional DRL-based controls, the proposed approach is smarter in that we can achieve a high level of generalization ability with a more compact neural network architecture and a smaller amount of training data. The controller is verified via practical application to active damping for a complex powertrain system with nonlinearities and parametric variations. Comparative tests demonstrate the high robustness of the proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。