提出RED框架,让平均奖励MDP实现多任务学习与风险感知。
Burning RED: Unlocking Subtask-Driven Reinforcement Learning and Risk-Awareness in Average-Reward Markov Decision Processes
- 利用平均奖励MDP的结构特性设计新算法框架
- 首次在线优化条件风险价值(CVaR)指标
- 无需双层优化或扩展状态空间,适合复杂决策场景
平均奖励马尔可夫决策过程(Average-reward MDPs)是不确定性下序列决策的基础框架,但其在强化学习(RL)中长期未受重视,多数研究集中于折扣型MDPs。本文揭示了平均奖励MDP的独特结构性质,并据此提出奖励扩展差分(Reward-Extended Differential, RED)强化学习框架,可在平均奖励设定下高效、有效地同时解决多种学习目标或子任务。我们构建了一套用于预测与控制的RED算法家族,包含表格情形下的收敛性证明算法。进一步展示该框架能力:首次实现完全在线条件下对经典的条件风险价值(CVaR)风险度量的优化,无需显式双层优化或扩展状态空间。
原文摘要 · Abstract (English)
Average-reward Markov decision processes (MDPs) provide a foundational framework for sequential decision-making under uncertainty. However, average-reward MDPs have remained largely unexplored in reinforcement learning (RL) settings, with the majority of RL-based efforts having been allocated to discounted MDPs. In this work, we study a unique structural property of average-reward MDPs and utilize it to introduce Reward-Extended Differential (or RED) reinforcement learning: a novel RL framework that can be used to effectively and efficiently solve various learning objectives, or subtasks, simultaneously in the average-reward setting. We introduce a family of RED learning algorithms for prediction and control, including proven-convergent algorithms for the tabular case. We then showcase the power of these algorithms by demonstrating how they can be used to learn a policy that optimizes, for the first time, the well-known conditional value-at-risk (CVaR) risk measure in a fully-online manner, without the use of an explicit bi-level optimization scheme or an augmented state-space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。