为强化学习设计信息论下界,确保策略在未知环境中的鲁棒性。
Information-Theoretic Minimax Regret Bounds for Reinforcement Learning based on Duality
- 基于对偶原理构建马尔可夫决策过程的极小极大后悔上界
- 首次给出有限时域下的信息论型极小极大后悔界
- 适用于需鲁棒决策的强化学习场景,如安全控制
研究智能体在未知环境中的行为,目标是寻找在所有可能环境中均能获得高累计回报的鲁棒策略。为此,考虑最小化不同环境参数下的最大后悔值,即极小极大后悔。本文聚焦于有限时域马尔可夫决策过程(MDPs)中极小极大后悔的信息论界。借鉴监督学习中的最小超额风险(MER)与极小极大超额风险概念,结合贝叶斯后悔的最新界,推导出极小极大后悔界。具体而言,建立了极小极大定理,并利用贝叶斯后悔界进行极小极大后悔分析。主要贡献包括:在MDP背景下定义合适的极小极大后悔形式,建立其信息论界,并将其应用于多种场景。
原文摘要 · Abstract (English)
We study agents acting in an unknown environment where the agent's goal is to find a robust policy. We consider robust policies as policies that achieve high cumulative rewards for all possible environments. To this end, we consider agents minimizing the maximum regret over different environment parameters, leading to the study of minimax regret. This research focuses on deriving information-theoretic bounds for minimax regret in Markov Decision Processes (MDPs) with a finite time horizon. Building on concepts from supervised learning, such as minimum excess risk (MER) and minimax excess risk, we use recent bounds on the Bayesian regret to derive minimax regret bounds. Specifically, we establish minimax theorems and use bounds on the Bayesian regret to perform minimax regret analysis using these minimax theorems. Our contributions include defining a suitable minimax regret in the context of MDPs, finding information-theoretic bounds for it, and applying these bounds in various scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。