提出在线算法评估动态风险下的强化学习策略,适用于无模拟器的真实场景。
Online Policy Evaluation for MDPs with Dynamic UBSR Measures

- 基于线性函数近似设计在线更新的UBSR-TD算法
- 理论证明算法几乎必然收敛,且在库存管理中表现优异
- 适合关注实时风险控制的研究者与工业应用者
在风险感知强化学习中,高效函数逼近方法是核心挑战。现有方法或局限于特定风险度量类别,或依赖模拟器,难以应用于完全在线场景。本文针对具有动态效用基短缺风险(UBSR)度量的马尔可夫决策过程(MDPs),提出计算高效的在线学习算法,并引入UBSR-TD算法,在线性函数近似下建立其几乎必然收敛的条件,同时设计多个加速收敛的变体。我们的框架表明,通过将损失函数融入时序差分误差,可直接将原有风险中性MDP的评估算法扩展至动态UBSR设置。数值实验验证了理论结果,且在具有保质期不确定性的易腐库存管理问题上的应用展示了方法的实际有效性。
原文摘要 · Abstract (English)
Developing efficient function-approximation methods for policy evaluation is a fundamental challenge in risk-aware reinforcement learning. Existing approaches either focus on restrictive classes of risk measures or rely on access to a simulator, limiting their applicability in fully online settings. In this work, we propose computationally efficient online learning algorithms for policy evaluation in Markov decision processes (MDPs) with dynamic utility-based shortfall risk (UBSR) measures under linear function approximation. Specifically, we introduce the UBSR-TD algorithm, establish conditions under which it converges almost surely, and develop several variants designed to accelerate convergence. Our formulation shows that existing policy evaluation algorithms for risk-neutral MDPs can be readily adapted to dynamic UBSR settings by incorporating a loss function into the temporal-difference error. Numerical experiments support our theoretical findings, and an application to a perishable inventory management problem with shelf-life uncertainty demonstrates the practical effectiveness of the proposed methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。