用少量数据快速适应不确定非线性系统,提升跟踪控制性能
Meta-Learning for Rapid Adaptation in Reference Tracking of Uncertain Nonlinear Systems

- 基于元学习框架,利用相似系统数据预训练共享动态表征
- 仅需少量目标系统数据和有限迭代步数即可实现高效控制优化
- 适用于数据稀缺场景,适合机器人、自动化等实时控制应用
本文研究不确定非线性系统的参考轨迹跟踪问题。由于从目标系统采集数据往往困难,我们的目标是仅使用有限的目标系统数据设计最优控制器。元学习通过利用与目标系统结构相似的源系统离线数据,为加速训练和提升控制性能提供了可行路径。受此启发,我们提出一种面向控制任务的元学习框架,将隐式模型无关元学习(iMAML)算法适配至控制场景。该框架分为两个阶段:离线元训练阶段,从源系统数据中学习一个聚合表示以捕捉相似系统的共享动态;在线元自适应阶段,仅使用少量目标系统数据和有限适应步数对该表示进行微调。我们将该框架建模为双层优化问题,并提出一种存储复杂度低、近似少的高效求解方法。所提框架具有通用性,可集成多种学习算法。为验证灵活性,我们分别基于神经状态空间模型和深度Q网络提出了两种具体算法,其主要区别在于是否需要显式系统辨识。数值仿真与硬件实验表明,该方法显著提升控制性能,持续优于基线方法。
原文摘要 · Abstract (English)
In this paper, we address the problem of reference tracking for uncertain nonlinear systems. Since collecting data from the target system (i.e., the system of interest) is often challenging, our objective is to design optimal controllers using limited target system data. Meta-learning provides a promising paradigm by leveraging offline data from source systems (systems sharing structural similarities with the target system) to accelerate training and enhance control performance. Motivated by this idea, we propose a meta-learning-based control framework that tailors the implicit model-agnostic meta-learning (iMAML) algorithm to the control setting. The framework operates in two phases: an (offline) meta-training phase, where an aggregated representation is learned from source data to capture the shared system dynamics among similar systems, and an (online) meta-adaptation phase, where this representation is fine-tuned on the target system using only a few data samples and limited adaptation steps. We formulate this framework as a bi-level optimization problem and provide an efficient solution with reduced storage complexity and few approximations. The proposed framework is general, allowing various learning algorithms to be integrated. To demonstrate this flexibility, we propose two specific learning algorithms that can be incorporated into our framework based on a neural state-space model and a deep Q-network, respectively. The primary distinction between these approaches is whether explicit system identification is required. Numerical simulations and hardware experiments demonstrate that the proposed methods enhance control performance and consistently outperform baseline approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。