提出新方法提升未知动态线性系统控制效率
The Confusing Instance Principle for Online Linear Quadratic Control
- 基于混淆实例原则设计新型控制算法
- 在多种场景下表现优于传统方法
- 适合大规模强化学习控制应用
我们重新审视了在未知动态下基于模型的强化学习对线性系统进行二次代价控制的问题。传统方法如面对不确定性时的乐观性(OFU)和汤普森采样,源自多臂赌博机(MABs),存在实际局限性。相比之下,我们提出一种基于混淆实例(CI)原则的新方法,该原则是MABs和离散马尔可夫决策过程(MDPs)中后悔下界的核心,并构成最小经验发散(MED)类算法的基础,后者在多种设置中表现出渐近最优性。通过结合LQR策略的结构以及灵敏度与稳定性分析,我们开发了MED-LQ。这一新型控制策略将CI和MED原则扩展至小规模之外。在综合性控制基准测试中,MED-LQ在多种场景下展现出竞争力,同时凸显其在大规模MDPs中的广泛应用潜力。
原文摘要 · Abstract (English)
We revisit the problem of controlling linear systems with quadratic cost under unknown dynamics with model-based reinforcement learning. Traditional methods like Optimism in the Face of Uncertainty and Thompson Sampling, rooted in multi-armed bandits (MABs), face practical limitations. In contrast, we propose an alternative based on the Confusing Instance (CI) principle, which underpins regret lower bounds in MABs and discrete Markov Decision Processes (MDPs) and is central to the Minimum Empirical Divergence (MED) family of algorithms, known for their asymptotic optimality in various settings. By leveraging the structure of LQR policies along with sensitivity and stability analysis, we develop MED-LQ. This novel control strategy extends the principles of CI and MED beyond small-scale settings. Our benchmarks on a comprehensive control suite demonstrate that MED-LQ achieves competitive performance in various scenarios while highlighting its potential for broader applications in large-scale MDPs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。