通过聚类相似控制系统,实现高效个性化控制策略学习。
Harnessing Data from Clustered LQR Systems: Personalized and Collaborative Policy Optimization
- 结合序列剔除与零阶优化,同时完成聚类与策略学习。
- 聚类准确率高,策略误差随聚类规模反比下降。
- 适合分布式多智能体系统,通信开销低,适合实际部署。
强化学习(RL)通常需要大量数据。为提升样本效率,可利用'近似相似'过程的数据。但因过程模型未知,难以判断哪些过程相似。本文在基准线性二次调节器(LQR)设置下研究此问题,考虑多个智能体,每个对应一个待控线性过程,其局部动态和任务可按相似性划分为簇。结合序列剔除与零阶策略优化思想,提出新算法,同步完成聚类与策略学习,为每簇输出个性化策略。在捕捉闭环性能差异的簇分离条件下,证明该方法以高概率正确聚类。进一步表明,各簇所学策略的次优性差距与簇大小成反比,无额外偏差,优于以往协作学习控制工作。本工作首次揭示聚类可用于数据驱动控制,实现协同统计增益而不受异质数据影响。从分布式实现角度看,方法仅需温和对数级通信开销。
原文摘要 · Abstract (English)
It is known that reinforcement learning (RL) is data-hungry. To improve sample-efficiency of RL, it has been proposed that the learning algorithm utilize data from 'approximately similar' processes. However, since the process models are unknown, identifying which other processes are similar poses a challenge. In this work, we study this problem in the context of the benchmark Linear Quadratic Regulator (LQR) setting. Specifically, we consider a setting with multiple agents, each corresponding to a copy of a linear process to be controlled. The agents' local processes can be partitioned into clusters based on similarities in dynamics and tasks. Combining ideas from sequential elimination and zeroth-order policy optimization, we propose a new algorithm that performs simultaneous clustering and learning to output a personalized policy (controller) for each cluster. Under a suitable notion of cluster separation that captures differences in closed-loop performance across systems, we prove that our approach guarantees correct clustering with high probability. Furthermore, we show that the sub-optimality gap of the policy learned for each cluster scales inversely with the size of the cluster, with no additional bias, unlike in prior works on collaborative learning-based control. Our work is the first to reveal how clustering can be used in data-driven control to learn personalized policies that enjoy statistical gains from collaboration but do not suffer sub-optimality due to inclusion of data from dissimilar processes. From a distributed implementation perspective, our method is attractive as it incurs only a mild logarithmic communication overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。