用Transformer学一个通用控制器,能适配多种线性系统。
Transformers As Generalizable Optimal Controllers
- 用Transformer从不同维度的系统中学习统一控制策略。
- 在多数系统上接近最优控制,且对参数扰动保持稳定。
- 只需轻量微调即可用于新系统,适合工程部署。
我们研究是否可由单一学习型控制器捕捉一组异构多输入多输出(MIMO)线性时不变(LTI)系统的最优状态反馈律。通过在不同状态和输入维度的系统上使用LQR生成轨迹训练Transformer策略,采用共享表征、标准化、填充、维度编码和掩码损失。该策略将近期状态历史映射为控制动作,推理时无需已知系统矩阵。在广泛系统上,其表现接近线性二次调节器(LQR),子优性极小,对中等参数扰动仍能保持稳定性,并可在未见系统上通过轻量微调提升性能。结果表明,Transformer策略是结构化线性系统族中近最优反馈律的实用近似。
原文摘要 · Abstract (English)
We study whether optimal state-feedback laws for a family of heterogeneous Multiple-Input, Multiple-Output (MIMO) Linear Time-Invariant (LTI) systems can be captured by a single learned controller. We train one transformer policy on LQR-generated trajectories from systems with different state and input dimensions, using a shared representation with standardization, padding, dimension encoding, and masked loss. The policy maps recent state history to control actions without requiring plant matrices at inference time. Across a broad set of systems, it achieves empirically small sub-optimality relative to Linear Quadratic Regulator (LQR), remains stabilizing under moderate parameter perturbations, and benefits from lightweight fine-tuning on unseen systems. These results support transformer policies as practical approximators of near-optimal feedback laws over structured linear-system families.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。