无需连续性假设,即可实现风险敏感强化学习的最优后悔率。
Continuity-Free Near-Minimax Leading-Order Regret for CVaR-UCBVI
- 采用自约束预算机制,动态控制价值估计误差。
- 在任意回报分布下达到近似最优的 $ ilde{O}( au^{-1/2} oot{2}{SAK})$ 领先后悔率。
- 适用于原子、混合与连续分布,适合高风险决策场景研究者。
针对有限时域表格型风险价值(CVaR)强化学习,已有工作证明:在任意归一化回报分布下,后悔率上界为 $ ilde{O}( au^{-1} oot{2}{SAK})$;在密度下界假设下可达到更优的 $ ilde{O}( oot{2}{SAK/ au})$。本文证明,相同的伯恩斯坦型 CVaR-UCBVI 算法在无连续性假设下仍可达成更优速率。关键在于引入选择性预算自约束:每轮回报不足的条件方差不超过 $ au$ 加上价值估计宽度。代入原始伯恩斯坦分解后,在高概率下得到任意归一化回报分布下的后悔率上界为 $ ilde{O}( oot{2}{SAK/ au} + (SAHK^{1/4} + S^2AH)/ au)$。其中 $ au^{-1/2}$ 主项与期望后悔率的极小极大下界一致,仅差对数因子。因此,该算法在全回报分布类中达到了领先阶次的极小极大最优性,低阶项仍保持 $ au^{-1}$ 依赖关系。
原文摘要 · Abstract (English)
For finite-horizon tabular CVaR reinforcement learning, prior work proves a $\widetilde{O}(\tau^{-1}\sqrt{SAK})$ leading regret bound for arbitrary normalized return laws and the sharper $\widetilde{O}(\sqrt{SAK/\tau})$ rate under a density lower bound. We show that the same Bernstein CVaR-UCBVI algorithm attains the sharper rate without continuity assumptions. The key is a selected-budget self-bound: the conditional variance of the episode shortfall is at most $\tau$ plus the value-estimation width. Substitution into the original Bernstein decomposition yields, with high probability, $\widetilde{O}(\sqrt{SAK/\tau}+(SAHK^{1/4}+S^2AH)/\tau)$ regret for arbitrary normalized return laws, including atomic, mixed, and continuous laws. The $\tau^{-1/2}$ leading term matches the expected-regret minimax lower bound up to logarithmic factors. Thus Bernstein CVaR-UCBVI is minimax-optimal over the full return-law class in the leading-order regime; the lower-order terms retain their $\tau^{-1}$ dependence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。