利用前瞻预测降低非平稳马尔可夫决策过程的累计损失
Predictive Control and Regret Analysis of Non-Stationary MDP with Look-ahead Information
- 基于前瞻预测设计低损失策略,动态调整行动
- 预测窗口越大,损失呈指数下降,误差可控
- 适合能源系统等有可靠预测的动态环境
非平稳马尔可夫决策过程(MDP)中的策略设计因系统转移和奖励随时间变化而极具挑战性,导致学习者难以确定最大化累积未来收益的最优动作。然而,在许多实际应用中,如能源系统,可获得前瞻预测信息,包括可再生能源发电和需求的预报。本文利用这些前瞻预测,提出一种算法,通过融合预测信息在非平稳MDP中实现低遗憾。理论分析表明,在一定假设下,遗憾随前瞻窗口扩大呈指数级下降。当系统预测存在误差时,只要误差随预测时长以亚指数方式增长,遗憾不会爆炸式上升。通过仿真验证了该方法在非平稳环境中的有效性。
原文摘要 · Abstract (English)
Policy design in non-stationary Markov Decision Processes (MDPs) is inherently challenging due to the complexities introduced by time-varying system transition and reward, which make it difficult for learners to determine the optimal actions for maximizing cumulative future rewards. Fortunately, in many practical applications, such as energy systems, look-ahead predictions are available, including forecasts for renewable energy generation and demand. In this paper, we leverage these look-ahead predictions and propose an algorithm designed to achieve low regret in non-stationary MDPs by incorporating such predictions. Our theoretical analysis demonstrates that, under certain assumptions, the regret decreases exponentially as the look-ahead window expands. When the system prediction is subject to error, the regret does not explode even if the prediction error grows sub-exponentially as a function of the prediction horizon. We validate our approach through simulations, confirming the efficacy of our algorithm in non-stationary environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。