用半马尔可夫决策过程提升模型预测控制的信息量,降低采样成本。
Increasing Information for Model Predictive Control with Semi-Markov Decision Processes
- 引入半马尔可夫框架实现时间抽象,扩展探索自由度。
- 相同采样预算下,数据信息量提升,样本复杂度下降。
- 适合强化学习中需高效探索的动态系统控制任务。
基于学习的动态系统模型预测控制近期工作通过信息论准则显著提升了样本效率。然而,系统的局部状态限制了序列探索机会,导致当前探索轨迹的观测信息量受限。本文通过引入半马尔可夫决策过程(Semi-Markov Decision Processes)实现时间抽象,突破该限制,提升固定采样预算下的总信息量,从而降低样本复杂度。
原文摘要 · Abstract (English)
Recent works in Learning-Based Model Predictive Control of dynamical systems show impressive sample complexity performances using criteria from Information Theory to accelerate the learning procedure. However, the sequential exploration opportunities are limited by the system local state, restraining the amount of information of the observations from the current exploration trajectory. This article resolves this limitation by introducing temporal abstraction through the framework of Semi-Markov Decision Processes. The framework increases the total information of the gathered data for a fixed sampling budget, thus reducing the sample complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。