arXiv:2510.01269cs.LGcs.SY2025-10中稿 · presentation at IC…

用LQR引导强化学习,安全控制结构振动

Safe Reinforcement Learning-Based Vibration Control: Overcoming Training Risks with LQR Guidance

  • 用随机模型生成的LQR策略引导RL训练,降低探索风险
  • 即使模型错误,LQR仍比无控状态更优,验证了引导有效性
  • 无需真实系统模型,适合工业界安全部署振动控制系统

外部激励引起的结构振动带来安全隐患、结构损伤和维护成本上升等风险。传统基于模型的控制方法如线性二次调节器(LQR)虽能有效抑制振动,但依赖精确系统模型,需繁琐的系统辨识。为避免此过程,可采用无需显式模型的强化学习(RL)方法,其仅通过观测结构行为学习控制策略。然而,若在真实结构上直接训练RL控制器,因缺乏先验知识而随机施加控制力,可能造成结构损伤。为此,本文提出以LQR控制器引导RL训练。我们发现,即使使用完全错误的模型,基于该模型的LQR策略仍优于无控制状态。据此,构建一种融合LQR与RL的混合控制框架:其中LQR策略由随机选取的模型参数生成,不依赖真实或近似系统模型,整体保持模型自由。该方法在不依赖显式系统模型的前提下,显著降低了原始RL训练中的探索风险。据我们所知,这是首个针对基于强化学习的振动控制训练安全问题提出并验证解决方案的研究。

原文摘要 · Abstract (English)

Structural vibrations induced by external excitations pose significant risks, including safety hazards for occupants, structural damage, and increased maintenance costs. While conventional model-based control strategies, such as Linear Quadratic Regulator (LQR), effectively mitigate vibrations, their reliance on accurate system models necessitates tedious system identification. This tedious system identification process can be avoided by using a model-free Reinforcement learning (RL) method. RL controllers derive their policies solely from observed structural behaviour, eliminating the requirement for an explicit structural model. For an RL controller to be truly model-free, its training must occur on the actual physical system rather than in simulation. However, during this training phase, the RL controller lacks prior knowledge and it exerts control force on the structure randomly, which can potentially harm the structure. To mitigate this risk, we propose guiding the RL controller using a Linear Quadratic Regulator (LQR) controller. While LQR control typically relies on an accurate structural model for optimal performance, our observations indicate that even an LQR controller based on an entirely incorrect model outperforms the uncontrolled scenario. Motivated by this finding, we introduce a hybrid control framework that integrates both LQR and RL controllers. In this approach, the LQR policy is derived from a randomly selected model and its parameters. As this LQR policy does not require knowledge of the true or an approximate structural model the overall framework remains model-free. This hybrid approach eliminates dependency on explicit system models while minimizing exploration risks inherent in naive RL implementations. As per our knowledge, this is the first study to address the critical training safety challenge of RL-based vibration control and provide a validated solution.

强化学习振动控制安全控制混合控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。