arXiv:2410.06828cs.LG2024-10被引 1

用简化模型训练控制策略,再迁移到完整系统,保证稳定性和效率。

Transfer Learning for a Class of Cascade Dynamical Systems

  • 在简化模型中训练策略,忽略部分状态动态,将其视为输入
  • 理论证明迁移效果依赖于内环控制器的稳定性,可量化保证
  • 适用于复杂系统控制,如四旋翼飞行器,兼顾计算效率与性能

本文研究强化学习中的迁移学习问题,针对在全状态系统中仿真耗时过长的情况,提出先在降阶系统中训练策略,再部署到全状态系统的方法。研究对象为一类级联动力系统,其中部分状态变量影响其余状态但反向不成立。策略在简化模型中训练时,忽略这些状态的动态,将其视为外部输入;在真实系统中则由经典控制器(如PID)处理。该结构使我们能基于内环控制器的稳定性,给出迁移效果的理论保证。数值实验以四旋翼系统验证了理论结果的有效性。

原文摘要 · Abstract (English)

This work considers the problem of transfer learning in the context of reinforcement learning. Specifically, we consider training a policy in a reduced order system and deploying it in the full state system. The motivation for this training strategy is that running simulations in the full-state system may take excessive time if the dynamics are complex. While transfer learning alleviates the computational issue, the transfer guarantees depend on the discrepancy between the two systems. In this work, we consider a class of cascade dynamical systems, where the dynamics of a subset of the state-space influence the rest of the states but not vice-versa. The reinforcement learning policy learns in a model that ignores the dynamics of these states and treats them as commanded inputs. In the full-state system, these dynamics are handled using a classic controller (e.g., a PID). These systems have vast applications in the control literature and their structure allows us to provide transfer guarantees that depend on the stability of the inner loop controller. Numerical experiments on a quadrotor support the theoretical findings.

强化学习迁移学习控制系统级联系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。